Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published 16 days ago • 10
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning Paper • 2607.28478 • Published 8 days ago • 4
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space Paper • 2608.01397 • Published 5 days ago • 5
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Paper • 2607.27951 • Published 8 days ago • 6
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Paper • 2607.27888 • Published 8 days ago • 8
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Paper • 2608.03874 • Published 3 days ago • 11
ExplainBench: Evaluating Code Explanations from Agents Paper • 2607.26451 • Published 9 days ago • 12
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems Paper • 2607.29241 • Published 7 days ago • 10
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Paper • 2608.00155 • Published 7 days ago • 12
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 3 days ago • 28
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation Paper • 2607.29209 • Published 7 days ago • 34
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 7 days ago • 36
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published 8 days ago • 36
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published 4 days ago • 37
Meshy T2: Fast Native Mesh Generation with Flow Matching Paper • 2607.28675 • Published 10 days ago • 53
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis Paper • 2608.02437 • Published 4 days ago • 59