Laguna S 2.1 Collection Our most capable model to date, designed for long-horizon work. • 12 items • Updated 5 days ago • 28
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Paper • 2607.11683 • Published 14 days ago • 144
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 13 days ago • 226
view article Article Welcome Inkling by Thinking Machines +2 burtenshaw, merve, pcuenq, ariG23498 • 12 days ago • 124
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 18 days ago • 75
Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published about 1 month ago • 149
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 64
view article Article Agentic Resource Discovery: Let agents search burtenshaw, evalstate • Jun 17 • 20
view article Article Beyond LoRA: Can you beat the most popular fine-tuning technique? +2 BenjaminB, sayakpaul, hubnemo, kashif • Jun 18 • 83
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Paper • 2606.06036 • Published Jun 4 • 76
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 122
OCC-RAG: Optimal Cognitive Core for Faithful Question Answering Paper • 2606.00683 • Published May 30 • 102
view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 134
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 124