PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 7 days ago • 33
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning Paper • 2603.09692 • Published Mar 10 • 5
Learning to Trigger: Reinforcement Learning at the Large Hadron Collider Paper • 2606.23993 • Published Jun 27 • 6
view article Article Jupyter Agents: training LLMs to reason with notebooks +1 baptistecolle, hannayukhymenko, lvwerra • Sep 10, 2025 • 67