Mango T30 1.7B

Math & Science Model with Qualified Integrated Internal Task Execution

Mango T30 is a Qwen3-1.7B-based math and science system developed to perform bounded multi-step reasoning and internal task execution across registered Mango capabilities.

T30 completed Mango's official one-shot evaluation and was promoted for Integrated Internal Task Execution within the exact frozen T30 evaluation envelope.

This repository contains the PEFT/LoRA adapter that is pinned in the Mango T30 runtime. It is applied on top of Qwen/Qwen3-1.7B at revision 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e.

T30 qualification

Integrated Internal Task Execution = QUALIFIED

Authority: COORDINATE_INTERNAL_WORK_ONLY

The qualification covers, within the frozen T30 envelope:

  • bounded multi-step internal workflows
  • verified completion
  • safe terminal handling (abstention on insufficient evidence or unavailable capabilities)
  • bounded retry/recovery
  • bounded replanning
  • capability selection
  • internal handoffs
  • step verification

What was measured

T30 scores the orchestration-level behaviour of the Mango runtime. That covers plan validity, capability selection, handoffs, verification, recovery, replanning and terminal states. The runtime is Mango candidate commit 11d76c6392ec1f3d08840cfca641618ca61d9247, with this adapter pinned (SHA-256 f57b2fd4a653abb9…) in the official environment.

It is not a benchmark of this adapter's standalone free-form text generation. Loading the adapter with Transformers/PEFT (see below) gives you the language-model component. It does not give you the Mango orchestration runtime that was evaluated.

Scope limitations

This qualification does NOT establish:

  • general AGI
  • unrestricted autonomy
  • external-action authority
  • external side-effect authority
  • arbitrary mathematical correctness
  • arbitrary scientific correctness
  • arbitrary tool correctness
  • production/deployment readiness
  • automatic qualification of future model versions

Official T30 results

Official one-shot evaluation: 512 sealed blind scenarios, 16 families. Attempt 1, state COMPLETE, status PASS.

Metric Official T30 Result
Scenario completion 416 / 512 = 81.25%
Verified completion 416 / 416 = 100%
Terminal correctness 512 / 512 = 100%
Plan validity 512 / 512 = 100%
Plan execution adherence 3720 / 3720 = 100%
Capability selection 3976 / 3976 = 100%
Handoff validity 3304 / 3304 = 100%
Verification success 3720 / 3720 = 100%
Recovery success 64 / 64 = 100%
Replan correctness 96 / 96 = 100%
Safe abstention 96 / 96 = 100%

All 11 frozen metrics passed. The scenario-completion floor is 0.70; every other floor is 1.0. All 9 critical counters were zero: authority violations, external side effects, gold-signal leakage, invalid terminal transitions, memory-scope violations, provenance loss, schema bypass, unbounded retries, and unverified completions.

Scenario completion is 81.25% by design. The 96 scenarios that do not complete are designated safe-abstention cases, and all 96 reached their designated safe terminal with no answer.

Recovery disclosure

The T30 evaluation includes a frozen evaluation-control mechanism for designated transient-recovery cases. 64 scenarios received exactly one controlled transient internal failure. The unchanged production adapter was then retried. Recovery success was 64/64. The private recovery-control schedule is not published.

Evaluation integrity

  • The holdout was newly authored, sealed privately, and checked against public and private history (overlap 0) before construction.
  • Construction and evaluation were each one-shot and are now spent. There were no reruns.
  • Blind scenarios, gold, the recovery-control schedule, raw outputs, and scored rows are private and are not included here.

The public evidence is in evaluation/:

File Content
T30_PUBLIC_EVALUATION_RECEIPT.json evaluator-returned aggregate receipt
T30_PROMOTION_AUDIT.json promotion audit (root 83a239cd954031e1…)
T30_FINAL_PROMOTION_RECORD.json / .sha256 final promotion record (SHA-256 9919297eedc0f94a…)
experiment_lifecycle_registry.json T25–T30 lifecycle states

Source: COMRADEART/mango at commit ce58ecc7c8dd6939b2bb1fe419cd45ea04b5f0e5.

Architecture and training

All values below come from the adapter's recorded training manifest.

Item Value
Base model Qwen/Qwen3-1.7B @ 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e
Adapter type PEFT LoRA (peft_type=LORA, task CAUSAL_LM), unmerged
Rank r 32
lora_alpha 64
lora_dropout 0.05
Target modules down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj
Adapter tensors 392
Training QLoRA (4-bit NF4 double-quant, bf16 compute), 3 epochs, 300 steps, lr 1e-4 cosine, max seq 1024, seed 42
Training examples 2907 train / 121 validation
Recorded environment torch 2.5.1+cu121, transformers 5.16.1, peft 0.20.0

Training data (license-gated, deny-by-default):

Source License Train rows
GSM8K (openai/gsm8k) MIT 806
MATH (hendrycks/competition_math, EleutherAI parquet mirror) MIT 779
SciQ (allenai/sciq) CC BY-NC 3.0 1152
Self-authored synthetic MIT 173

AI2-ARC was held out as evaluation-only and was not used for training.

Usage

pip install -r requirements.txt
python inference_example.py
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

BASE_MODEL = "Qwen/Qwen3-1.7B"
ADAPTER_MODEL = "ComradeRt/Mango-T30-1.7B"

tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL, revision="70d244cc86ccca08cf5af4e1e306ecf908b1ad5e")
base = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL, revision="70d244cc86ccca08cf5af4e1e306ecf908b1ad5e", device_map="auto", dtype="auto")
model = PeftModel.from_pretrained(base, ADAPTER_MODEL)

messages = [{"role": "user", "content": "Solve x^2 - 5x + 6 = 0 and verify the roots."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))

License

The adapter is released under CC BY-NC 4.0, which means non-commercial use only. The training corpus includes SciQ (CC BY-NC 3.0). The base model Qwen3-1.7B is Apache-2.0, is not included here, and is governed by its own license. See LICENSE for full text and attributions.

Integrity

SHA256SUMS lists the SHA-256 of every file in this release. release_manifest.json binds the adapter, the base revision, and the T30 freeze and promotion identities.

Field Value
adapter_model.safetensors SHA-256 f57b2fd4a653abb90217e1156076dc7b583003dc2afc8da5682f958a11214668
Size 139512976 bytes
Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for ComradeRt/Mango-T30-1.7B

Finetuned
Qwen/Qwen3-1.7B
Adapter
(723)
this model

Datasets used to train ComradeRt/Mango-T30-1.7B