Request access to Fin-o1-8B

Fin-o1-8B is released by The Fin AI for research. Access is granted automatically after you complete this short form.

This model is released for research purposes. It must not be used to provide personalized investment advice or to make automated financial decisions without human oversight. Use is also subject to the license of the base model (Qwen3, Apache 2.0).

Log in or Sign Up to review the conditions and access this model content.

Fin-o1-8B

📄 Paper · 🤗 Collection · 💻 Code · 🏆 Leaderboard · 🌐 The Fin AI

Fin-o1-8B is a financial reasoning model fine-tuned from Qwen3-8B with supervised fine-tuning on financial chain-of-thought data followed by GRPO reinforcement learning. It was introduced in Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance (arXiv:2502.08127).

Fin-o1 vs. Fino1. The two families come from the same paper but are different generations:

Fino1 (2025-02/03) Fin-o1 (2025-05)
8B Fino1-8B — Llama-3.1-8B-Instruct Fin-o1-8B — Qwen3-8B
14B Fino1-14B — Qwen2.5-14B-Instruct Fin-o1-14B — Qwen3-14B
Training CoT SFT + RL SFT + GRPO

For new work we recommend the Fin-o1 models.

Model Details

Base model Qwen/Qwen3-8B
Architecture Qwen3ForCausalLM, 36 layers, hidden size 4096
Precision bfloat16
Context length 40,960 tokens (max_position_embeddings)
Training data TheFinAI/FinCoT — reasoning paths derived from FinQA, TAT-QA, DocMath-Eval, Econ-Logic, BizBench-QA and DocFinQA
Training method SFT, then GRPO with accuracy and format rewards
Language English
License Apache 2.0 (same as the base model)

GRPO stage (from trainer_state.json / train_results.json in this repo)

Steps / epochs 500 steps, 2 epochs
Training samples 1,525
Per-device batch size 2
Peak learning rate 3e-6
Rewards accuracy_reward, format_reward
Wall-clock ≈ 2.2 h

Quick Start

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "TheFinAI/Fin-o1-8B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")

messages = [{
    "role": "user",
    "content": "A company's revenue grew from $120M to $150M while operating costs rose from $90M to $105M. "
               "By how many percentage points did the operating margin change?",
}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to(model.device)
output = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Evaluation

Fin-o1 and Fino1 are evaluated on FinQA, DocMath-Eval (simple / complex long), XBRL-Math and related financial reasoning tasks. See the paper and the Open FinLLM Reasoning Leaderboard for results.

Intended Use & Limitations

  • Intended use: research on financial numerical reasoning, question answering over financial text and tables, and as a baseline for financial reasoning models.
  • Not intended for: investment advice, trading decisions, or any automated financial decision without human review.
  • The model can produce incorrect calculations or hallucinated figures, especially on long documents with multiple tables.
  • Trained and evaluated on English data only.

Citation

@misc{qian2025fino1transferabilityreasoningenhancedllms,
      title={Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance},
      author={Lingfei Qian and Weipeng Zhou and Yan Wang and Xueqing Peng and Han Yi and Yilun Zhao and Jimin Huang and Qianqian Xie and Jian-yun Nie},
      year={2025},
      eprint={2502.08127},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2502.08127},
}

Contact

Questions and issues: open a discussion on this repository or on GitHub.

Downloads last month
245
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
Input a message to start chatting with TheFinAI/Fin-o1-8B.

Model tree for TheFinAI/Fin-o1-8B

Finetuned
Qwen/Qwen3-8B
Finetuned
(2161)
this model
Finetunes
2 models
Merges
1 model
Quantizations
1 model

Dataset used to train TheFinAI/Fin-o1-8B

Collection including TheFinAI/Fin-o1-8B

Paper for TheFinAI/Fin-o1-8B