WHALE: A Simple Recipe for Joint Harness-Weight Optimization Paper • 2609.00196 • Published 15 days ago • 36
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 14 days ago • 121
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs Paper • 2609.04753 • Published 11 days ago • 17
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers Paper • 2604.07822 • Published Apr 9 • 3
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published Jul 3 • 85
Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models Paper • 2606.12689 • Published Jun 10 • 2
Retrieval-Augmented LLM Agents: Learning to Learn from Experience Paper • 2603.18272 • Published Mar 18 • 1
MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution Paper • 2603.18718 • Published Mar 19 • 10
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens Paper • 2603.23516 • Published Mar 6 • 53
ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation Paper • 2601.21420 • Published Jan 29 • 42
KaVa: Latent Reasoning via Compressed KV-Cache Distillation Paper • 2510.02312 • Published Oct 2, 2025 • 4
view article Article mem-agent: Equipping LLM Agents with Memory Using RL driaforall • Oct 9, 2025 • 33
view article Article Provence: efficient and robust context pruning for retrieval-augmented generation nadiinchi • Jan 28, 2025 • 26
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates Paper • 2502.06772 • Published Feb 10, 2025 • 22
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration Paper • 2502.01068 • Published Feb 3, 2025 • 18
mHuBERT-147 models Collection Compact yet powerful multilingual speech representation models based on the HuBERT architecture. • 3 items • Updated Jun 4, 2024 • 8