James4Ever0/computer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcp Updated Aug 10, 2025 • 153 • 6
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Paper • 2610.01892 • Published 10 days ago • 31
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 5 days ago • 88
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning Paper • 2610.07043 • Published 6 days ago • 42
Base Models Can Reason By Taking a Cue From Training Data Paper • 2610.06851 • Published 6 days ago • 24
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 6 days ago • 97
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points Paper • 2610.06833 • Published 6 days ago • 32
jaehyeokdoo2/openpi-droid-pnpcarrot-singetask-qflow-offlinerl-criticwarmup2000-alpha100-bs8-test Updated Mar 4 • 5
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 9 days ago • 33
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Paper • 2609.39071 • Published 11 days ago • 65