ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research Paper • 2606.07591 • Published May 28 • 102
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research Paper • 2606.09730 • Published Jun 8 • 55
DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch Paper • 2606.10728 • Published Jun 9 • 36
Towards Diverse Scientific Hypothesis Search with Large Language Models Paper • 2606.10587 • Published Jun 9 • 2
TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search Paper • 2606.11662 • Published Jun 10 • 10
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 31
Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences Paper • 2606.16905 • Published Jun 15 • 7
Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark Paper • 2606.18648 • Published Jun 17 • 15
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 66
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published Jun 26 • 11
AI translation of literary texts is "fine", but readers still prefer human translations Paper • 2606.26040 • Published Jun 24 • 9