Add TerminalBench 2.1 and SWE Atlas QnA evaluation results

#6

Adds the remaining coding evaluation results reported in the model card using the Hugging Face eval-results format.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment