LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
9/10LongStraw presents methods to enable reinforcement learning with extremely long contexts exceeding 2 million tokens while operating within fixed GPU memory budgets. It addresses the significant gap between training context length and inference context limitation, optimizing GPU usage to support ultra-long context LLM applications.
