R
RTPO
Reinforcement learning framework for stabilizing multi-turn agent training through reverse credit assignment
Open SourceFree
About
RTPO (Reverse-Turn Policy Optimization) is a research framework addressing training instability in multi-turn agentic reinforcement learning. It introduces novel credit assignment mechanisms that work backward through conversation turns, improving stability for agents handling complex reasoning workflows. Particularly relevant for training agents that require multi-step planning and contextual understanding across extended interactions.
Details
| Language | |
| Patterns |
Tags
frameworkmulti-agentautonomousopen-sourcepython
Quick Info
- Organization
- Research (Yugu Li et al.)
- Pricing
- open-source
- Free Tier
- Yes
- Updated
- Aug 20, 2026
Also in Frameworks
L
LangChain
Build context-aware reasoning applications with LLMs
OSSFree
LangChain AI
144.7K850.0K/wtoday467
A
AutoGen
Microsoft's framework for building multi-agent AI systems
OSSFree
Microsoft
60.6K18w ago444
C
CrewAI
Multi-agent orchestration framework for collaborative AI workflows
OSSFree
CrewAI Inc
57.4Ktoday301