AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents Paper • 2610.05140 • Published 8 days ago • 43
AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data Viewer • Updated Feb 1 • 5.28k • 270 • 14
DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation Paper • 2610.04596 • Published 9 days ago • 20
NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale Paper • 2610.08430 • Published 6 days ago • 24
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation Paper • 2610.02781 • Published 10 days ago • 16
HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models Paper • 2610.07727 • Published 6 days ago • 16
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Paper • 2610.01892 • Published 11 days ago • 33
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding Paper • 2610.07342 • Published 7 days ago • 21
Learning Functional Subspaces for Neural Network Compression Paper • 2609.40127 • Published 12 days ago • 25
andersonbcdefg/red_teaming_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 1, 2023 • 35.3k • 279 • 13
crumb/bloom-560m-RLHF-SD2-prompter-aesthetic Text Generation • 0.6B • Updated Mar 19, 2023 • 342 • 29
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 6 days ago • 92
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning Paper • 2610.07043 • Published 7 days ago • 42