Instructions to use eltonsarmanho/qwen3-1.7b-questoes-matematica with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M # Run inference directly in the terminal: llama cli -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M # Run inference directly in the terminal: llama cli -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Use Docker
docker model run hf.co/eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "eltonsarmanho/qwen3-1.7b-questoes-matematica" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "eltonsarmanho/qwen3-1.7b-questoes-matematica", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
- Ollama
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with Ollama:
ollama run hf.co/eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
- Unsloth Desktop
- Pi
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with Docker Model Runner:
docker model run hf.co/eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
- Lemonade
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Run and chat with the model
lemonade run user.qwen3-1.7b-questoes-matematica-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use eltonsarmanho/qwen3-1.7b-questoes-matematica with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "eltonsarmanho/qwen3-1.7b-questoes-matematica:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3-1.7B SAEB Math Question Generator (offline / on-device)
English
Summary
This model is a QLoRA fine-tune of Qwen3-1.7B that writes multiple-choice mathematics questions in the style of SAEB, Brazil's national assessment of basic education. Each output is a structured JSON object ready for a mobile app, and the model is packaged as a quantized GGUF so it can run fully offline on phones and tablets through llama.cpp.
We built it for a concrete classroom constraint: many Brazilian schools have unreliable connectivity, so teachers need a question generator that works on the device in their hands. A 1.7B-parameter model is small enough to fit there. Out of the box, though, it is slow and rarely follows the required format. This fine-tune is our attempt to fix both.
| Code and data | https://github.com/eltonsarmanho/TreinamentoModeloQuestoes |
| Base model | unsloth/Qwen3-1.7B (4-bit NF4) |
| Method | QLoRA, supervised fine-tuning with TRL SFTTrainer on Unsloth |
| Deployment format | GGUF, Q4_K_M quantization (~1.1 GB), llama.cpp |
| Output language | Brazilian Portuguese |
| Data cycle | L2-2026-09, promoted on 2026-09-16 (8/8 blocking gates passed) |
| Artifact hash | sha256 0e5ceb6ef929ee86…; previous model kept in baseline_v1/ (sha256 b461b10fbf23622d…) for rollback |
Repository contents
| Path | Contents |
|---|---|
lora/ |
Unmerged LoRA adapters, for re-merging or re-quantizing with another method |
gguf/ |
Merged model quantized to Q4_K_M, ready for llama.cpp |
data/ |
train.jsonl (963 examples: 719 real + 244 synthetic) and val.jsonl (30 real items, frozen across cycles) |
eval_report.json |
Structural quality, perplexity and speed metrics |
Motivation and design
The base model had two practical problems in the app: it spent many tokens on free-form "thinking" and often returned output the app could not parse. We handled them separately.
- Format and quality. We fine-tuned on real SAEB items, plus synthetic
arithmetic items whose answers are computed in Python. The model learns a
fixed JSON contract defined by the mobile team:
{"questoes": [...]}, where each question has a stem, five options (A–E), a step-by-step solution, the correct letter, and a difficulty label (EASY/MEDIUM/HARD). - Speed. Training and inference use non-thinking mode
(
enable_thinking=False). Outputs are short (about 150–300 tokens), and the model ships as GGUF Q4_K_M, the de facto format for offline LLMs on Android and iOS.
Training data
Current cycle: L2-2026-09
The original bank had 553 items. In this cycle, contributors submitted 499
new, original questions for grades 1 to 4, which the bank had not covered
before. Every item went through an automated check
(src/validar_entrega.py) and a manual semantic review.
| Submission outcome | Items |
|---|---|
| Submitted | 499 |
| Approved | 481 |
| Approved with warning | 5 |
| Rejected | 13 |
| Added to the bank | 486 |
All 13 rejections share one skill, EF03MA01 (grade 3). The questions asked who had more (for example, "WHO has more stickers?") but were answered with a quantity, and the step-by-step solution never carried out the comparison. The automated check missed this; the semantic review caught it, and the items went back to their authors.
The bank now has 553 + 486 = 1,039 items. Ten of the new items depend on
an image and are excluded from text-only training, so 476 reach the training
set. For the 170 questions whose figure was rewritten as text, the original
PNG is kept in a separate column (imagem_path). This avoids repeating an
earlier mistake in which 230 questions were lost from the bank.
Training filter (src/extract_data.py): mathematics only, fully textual
questions only, and no duplicate options.
| Previous baseline | Cycle L2-2026-09 |
|
|---|---|---|
| Valid real questions | 303 | 779 |
| Split | 273 train / 30 val | 719 train / 30 val |
| Grades covered | 2, 5, 9 | 1, 2, 3, 4, 5, 9 |
| Real items verifiable by machine | 19.8% | 68.8% |
Evaluation sets. data/val.jsonl (30 items) is frozen: it is identical,
item by item, to the previous cycle's set, so the two models can be compared
fairly on the same questions. None of the new items were added to it. A
second set, data/val_novos_v1.jsonl (30 items, 3 per new skill, grades
1–4), is held out from training to measure coverage of the new grades.
Synthetic augmentation (generate_synthetic.py, unchanged this cycle).
We add 244 synthetic examples (396 questions, since some examples bundle
2–5 questions in one "questoes": [...] list). They cover addition,
subtraction, multiplication, division, percentages, powers and fractions
(skills H07–H09). The answer is always computed in Python before the
question is written, never by an LLM. Each item reuses a real
(grade, skill, description, difficulty) tuple from the bank. The synthetic
block is generated from source code, not from the bank, so it is
byte-identical across cycles.
Final training set: 963 examples (719 real + 244 synthetic).
Example format (chat SFT, unchanged this cycle):
system: a fixed instruction describing the role and the JSON schema.user:"Gere {quantidade} questão(ões) de matemática. Ano: {ano}. Habilidade: {habilidade} — {descrição}. Dificuldade: {dificuldade}."assistant:{"questoes": [{"enunciado", "alternativas" (A–E), "resolucao_passo_a_passo", "resposta_correta", "difficulty"}, ...]}
The model writes resolucao_passo_a_passo (the solution) before
resposta_correta (the answer letter), so it reasons first and commits to an
answer second.
Training procedure
We use QLoRA (Dettmers et al., 2023) with Unsloth and TRL's SFTTrainer. The base model stays frozen in 4-bit NF4 and only the LoRA adapters are trained. This fits on a laptop GPU (RTX 3060 Laptop, 6 GB VRAM). We left the hyperparameters untouched this cycle; nothing in our experiments argued for changing them.
| Hyperparameter | Value |
|---|---|
| Base model | unsloth/Qwen3-1.7B (unsloth/qwen3-1.7b-unsloth-bnb-4bit) |
max_seq_length |
1024 |
| LoRA rank / alpha | 16 / 32 |
| LoRA dropout | 0.0 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Effective batch size | 2 × 8 gradient accumulation = 16 |
| Learning rate | 2e-4, cosine schedule, 5% warmup |
| Epochs | up to 3, early stopping on eval_loss (patience 2) |
| Optimizer | adamw_8bit, weight decay 0.01 |
| Precision | bf16 |
| Loss | assistant tokens only (train on completions) |
| Seed | 42 |
eval_loss per epoch: 0.5772 → 0.5362 (best, epoch 2) → 0.5414.
Final train_loss: 0.4424.
Framework versions: TRL 0.24.0, Transformers 5.5.0, PyTorch 2.10.0.
Controlling for the software update
Between cycles the original Python environment was lost and rebuilt with a newer stack (torch 2.10, transformers 5.5, unsloth 2026.9.4). To separate the effect of the new data from the effect of the new libraries, we trained a control model on the same 517 examples as the previous baseline, under the new stack, with the same code and hyperparameters. The dataset is the only variable.
| Metric (frozen set, n=30, seed-paired) | Control (old data) | Candidate (new data) |
|---|---|---|
eval_loss (best epoch) |
0.5583 | 0.5362 (−4.0%) |
| Mentions of the word "figura" | 10.0% | 10.0% (identical, so this comes from the stack, not the data) |
Evaluation
We evaluated on two sets, generating with the GGUF model in llama.cpp under a
GBNF grammar, temperature=0.7, top_p=0.8, and paired seeds per sample.
Frozen set (grades 2, 5, 9; n=30; same as previous cycle)
| Structural metric | Result |
|---|---|
| Valid JSON / valid wrapper / complete schema | 100% |
Valid resposta_correta (A–E) |
100% |
| Five distinct options | 100% |
Valid difficulty that matches the request |
100% |
| Answer–solution consistency¹ | 100% (7/30 samples verifiable) |
| Unresolved dependence on a missing visual² | 3.3% (1/30); not significant vs. baseline (McNemar p = 1.000) |
| Answer-key bias (most frequent letter) | 36.7% (baseline: 46.7%) |
| Unrecoverable failures after best-of-N | 0 |
¹ Consistency is a deterministic check (schema_utils.check_consistency,
no LLM involved). It extracts an arithmetic expression of the form
"a op b = r" from the solution and compares it with the option selected in
resposta_correta. It only applies when such an expression can be found
(7 of 30 samples here). Questions with verbal reasoning or textual or
geometric answers are not counted. In grade 9, coverage falls to 2/19, so
most grade-9 generations cannot be verified automatically. We treat this
as an open quality risk (see Limitations).
² depende_de_visual_ausente_pct replaces the older
mencoes_figura_pct as the decision metric. The old metric counted any
occurrence of words like "figure", "image", "graph" or "drawing", which gave
false positives such as "Drawing" listed as a hobby in an option, or a
self-contained question that merely asks students to build a bar chart.
The new metric (schema_utils.DEPENDENCIA_VISUAL_PATTERN) looks for
deictic references such as "look at the figure below" or "according to
the table". On our 14 test cases it made no mistakes, though 14 is a small sample. The
production pipeline (test_model.generate_validated) now also penalizes and
regenerates candidates that rely on a missing visual.
Coverage holdout (grades 1–4; n=30; excluded from training)
This is the main result of the cycle: performance on the grades that only the
L2-2026-09 batch introduced.
| Control (never saw these grades) | Candidate (trained on them) | |
|---|---|---|
| Machine-verifiable | 23/30 | 30/30 |
| Grade 1 | 9/9 | 9/9 |
| Grade 2 | 7/9 | 9/9 |
| Grade 3 | 4/9 | 9/9 |
| Grade 4 | 3/3 | 3/3 |
| Structure / difficulty / consistency | 100% | 100% |
| Dependence on a missing visual | 0% | 0% |
Paired McNemar test (7 wins, 0 losses): p = 0.016. This is the only statistically significant result of the cycle. The improvement on the frozen set (grades 2, 5, 9) is within noise (p = 0.625), which is what we expected: apart from grade 2, the new batch added no questions for those grades.
| Language metric | Result |
|---|---|
| Perplexity (reference answer) | not measured this cycle (gate G9, informational) |
| Speed (desktop CPU, GGUF, 4 threads) | Result |
|---|---|
| Generation throughput | 26.5 tokens/s (baseline: 26.3, same pipeline with the visual-dependence filter) |
| Mean total latency (includes model loading) | 17.9 s |
Promotion gates (src/promover_checkpoint.py)
| Gate | Criterion | Result |
|---|---|---|
| G1 Structural contract | ≥ 99% and ≥ baseline on 7 flags | 100% on all |
| G2 Failures after best-of-N | = 0 | 0 |
| G3 Consistency | ≥ baseline − 5 pp | 100% → 100% |
| G4 Visual dependence | no statistically significant increase | 0% → 3.3% (McNemar p = 1.000) |
| G5 Difficulty adherence | ≥ baseline − 5 pp | 100% → 100% |
| G6 Answer-key bias | ≤ baseline + 5 pp | 46.7% → 36.7% |
| G7 Forgetting (grades 5, 9) | no structural regression | preserved |
| G8 Speed | ≥ baseline − 10% | 26.3 → 26.5 tok/s |
| G9 Perplexity (informational) | — | not measured |
Decision: PROMOVIDO (promoted). All 8 blocking gates passed.
How to use
LoRA adapters with Unsloth
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="<this-repo>/lora",
max_seq_length=1024,
load_in_4bit=True,
)
FastLanguageModel.for_inference(model)
messages = [
{"role": "system", "content": "Você é um gerador de questões de matemática no padrão SAEB..."},
{"role": "user", "content": "Gere 1 questão(ões) de matemática. Ano: 3º ano. Habilidade: EF03MA01 — ... . Dificuldade: Fácil."},
]
inputs = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
enable_thinking=False, return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512, temperature=0.7, top_p=0.8)
print(tokenizer.decode(output[0, inputs.shape[1]:], skip_special_tokens=True))
GGUF with llama.cpp (offline / mobile)
llama-cli -m gguf/qwen3-1.7b.Q4_K_M.gguf --temp 0.7 --top-p 0.8 \
--grammar-file grammars/questao.gbnf \
-p "Gere 1 questão(ões) de matemática. Ano: 3º ano. Habilidade: EF03MA07 — ... . Dificuldade: Moderado."
We recommend always using the GBNF grammar (grammars/questao.gbnf). It
enforces the exact contract: the {"questoes": [...]} wrapper, five options
(A–E), resposta_correta in {A,B,C,D,E}, and difficulty in
{EASY,MEDIUM,HARD}. Combine it with generate_validated()
(src/test_model.py), which checks structure, consistency and visual
dependence before returning a question and regenerates when needed.
Limitations
- Mathematical verification is still weak for grade 9 (2/19 samples
verifiable this cycle). We observed one item with a mathematically wrong
solution: it computed
√(376.8/√3) ≈ 10√3when the correct value is about 14.7, and claimed that an equilateral triangle inscribed in a circle of radius 10 cm covers 30% of the circle's area (the true figure is about 41%). The pipeline accepted it with statusnao_verificavelbecause it had no recognizable "a op b = r" expression. This is the largest open quality risk. It was already present in the previous model, so it is not a regression, but this cycle did not fix it. A symbolic verifier (e.g.sympy) for grade-9 algebra and geometry skills is the natural next step. - The answer letter can disagree with the solution on samples the
deterministic checker cannot cover. This is expected from a 1.7B model
without explicit chain-of-thought. Best-of-N sampling plus deterministic
correction (
generate_validated()) reduces the problem but does not remove it. - The real data is concentrated. The
L2-2026-09batch makes up 61% of the real corpus but covers only 10 skills (about 48 items each), versus about 5.5 items for each of the 55 older skills. The model may over-specialize on those 10 skills; we monitor this withdata/val_novos_v1.jsonl. - The batch shifted the difficulty distribution. Easy items rose to 52% of the real corpus (from 27%) and hard items fell to 15% (from 35%). Difficulty adherence stayed at 100% on both sets, but the training distribution is now unbalanced and needs attention in future cycles.
- Option E in real bank items is always a fixed distractor ("None of the above"), because the original bank only has options A–D. Only the synthetic questions have five pedagogically distinct distractors.
- The training stack is not pinned.
requirements.txtdeliberately does not pin torch, unsloth or transformers, which made it impossible to reproduce the previous cycle's exact environment. We recommend freezing arequirements.lockfrom the stack validated here (torch 2.10.0, transformers 5.5.0, unsloth 2026.9.4, trl 0.24.0). - Not yet tested on real mobile hardware. All speed numbers come from a desktop CPU with 4 threads; on-device latency and memory use are unknown.
- Batch generation (
quantidade > 1questions per call) was not evaluated in this cycle, although the contract and the synthetic data support it.
Intended use and ethical considerations
The model is meant to support teachers, not replace them. Generated questions should be reviewed by an educator before reaching students, especially for grade 9 content, where automatic verification is limited. It is not designed for high-stakes assessment, and it does not generate questions that depend on images.
Data provenance
Questions come from an internal SAEB-style item bank used by the project
(553 original items) and from a batch of 486 original questions written by
contributors following TreinamentoDados/GUIA_ENTREGA_ALUNOS.md. The batch
was audited and incorporated in cycle L2-2026-09
(src/validar_entrega.py, src/importar_entregas.py). Training, extraction
and evaluation code are available in the GitHub repository linked above.
Citation
If you use this model, its data pipeline or its evaluation protocol, please cite:
@misc{sarmanho2026qwen3saeb,
author = {Sarmanho, Elton},
title = {{Qwen3-1.7B SAEB Math Question Generator}: On-Device
Generation of Structured Multiple-Choice Mathematics
Questions for Brazilian Basic Education},
year = {2026},
month = sep,
howpublished = {\url{https://github.com/eltonsarmanho/TreinamentoModeloQuestoes}},
note = {QLoRA fine-tune of Qwen3-1.7B, data cycle L2-2026-09,
GGUF Q4\_K\_M release}
}
Please also cite the underlying methods:
@inproceedings{dettmers2023qlora,
author = {Dettmers, Tim and Pagnoni, Artidoro and Holtzman, Ari and Zettlemoyer, Luke},
title = {{QLoRA}: Efficient Finetuning of Quantized {LLMs}},
booktitle = {Advances in Neural Information Processing Systems},
volume = {36},
year = {2023}
}
@misc{qwen3technicalreport,
author = {{Qwen Team}},
title = {{Qwen3} Technical Report},
year = {2025},
eprint = {2505.09388},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}
Português
Resumo
Este modelo é um ajuste fino (QLoRA) do Qwen3-1.7B que elabora questões de matemática de múltipla escolha no padrão do SAEB (Sistema de Avaliação da Educação Básica). Cada saída é um objeto JSON estruturado, pronto para o aplicativo mobile, e o modelo é distribuído em GGUF quantizado para rodar totalmente offline em celulares e tablets via llama.cpp.
O ponto de partida foi uma restrição real de sala de aula: muitas escolas brasileiras têm conexão instável, e o professor precisa de um gerador de questões que funcione no aparelho que tem em mãos. Um modelo de 1,7B de parâmetros cabe nesse aparelho, mas, sem ajuste, é lento e raramente segue o formato exigido. Este ajuste fino tenta resolver as duas coisas.
| Código e dados | https://github.com/eltonsarmanho/TreinamentoModeloQuestoes |
| Modelo base | unsloth/Qwen3-1.7B (4-bit NF4) |
| Método | QLoRA, ajuste supervisionado com TRL SFTTrainer sobre Unsloth |
| Formato de deploy | GGUF, quantização Q4_K_M (~1,1 GB), llama.cpp |
| Idioma de saída | Português do Brasil |
| Ciclo de dados | L2-2026-09, promovido em 16/09/2026 (8/8 gates bloqueantes aprovados) |
| Hash do artefato | sha256 0e5ceb6ef929ee86…; modelo anterior preservado em baseline_v1/ (sha256 b461b10fbf23622d…) para reversão |
Arquivos neste repositório
| Caminho | Conteúdo |
|---|---|
lora/ |
Adaptadores LoRA não mesclados, para re-mesclar ou re-quantizar com outro método |
gguf/ |
Modelo mesclado e quantizado em Q4_K_M, pronto para llama.cpp |
data/ |
train.jsonl (963 exemplos: 719 reais + 244 sintéticos) e val.jsonl (30 itens reais, congelado entre ciclos) |
eval_report.json |
Métricas de qualidade estrutural, perplexidade e velocidade |
Motivação e desenho
O modelo base tinha dois problemas práticos no app: gastava muitos tokens em "pensamento" livre e frequentemente devolvia uma saída que o app não conseguia interpretar. Tratamos cada um separadamente.
- Formato e qualidade. O ajuste fino usa questões reais do banco SAEB e
questões sintéticas de aritmética com resposta calculada em Python. O
modelo aprende um contrato JSON fixo, definido pela equipe do app:
{"questoes": [...]}, em que cada questão tem enunciado, cinco alternativas (A–E), resolução passo a passo, letra correta e dificuldade (EASY/MEDIUM/HARD). - Velocidade. Treino e inferência usam o modo non-thinking
(
enable_thinking=False). As saídas são curtas (cerca de 150–300 tokens) e o modelo é exportado em GGUF Q4_K_M, o formato de fato para LLMs offline em Android e iOS.
Dados de treinamento
Ciclo atual: L2-2026-09
O banco original tinha 553 itens. Neste ciclo, colaboradores entregaram 499
questões autorais para o 1º ao 4º ano, anos até então sem cobertura. Cada
item passou por uma checagem automática (src/validar_entrega.py) e por uma
revisão semântica manual.
| Resultado da auditoria | Itens |
|---|---|
| Entregues | 499 |
| Aprovados | 481 |
| Aprovados com alerta | 5 |
| Rejeitados | 13 |
| Incorporados ao banco | 486 |
As 13 rejeições têm a mesma habilidade, EF03MA01 (3º ano). As perguntas eram de identificação ("QUEM tem mais figurinhas?"), mas a resposta era uma quantidade, e a resolução passo a passo nunca fazia a comparação pedida. A checagem automática não pegou o defeito; a revisão semântica pegou, e os itens voltaram aos autores.
O banco passou a ter 553 + 486 = 1.039 itens. Dez dos itens novos
dependem de imagem e ficam fora do treino textual, então 476 chegam ao
treino. Nas 170 questões cuja figura foi reescrita em texto, o PNG original
é preservado numa coluna separada (imagem_path). Isso evita repetir um erro
anterior, em que 230 questões se perderam do banco.
Filtro de treino (src/extract_data.py): apenas matemática, apenas
questões 100% textuais e sem alternativas duplicadas.
| Baseline anterior | Ciclo L2-2026-09 |
|
|---|---|---|
| Questões reais válidas | 303 | 779 |
| Split | 273 treino / 30 val | 719 treino / 30 val |
| Anos cobertos | 2º, 5º, 9º | 1º, 2º, 3º, 4º, 5º, 9º |
| Itens reais verificáveis por máquina | 19,8% | 68,8% |
Conjuntos de avaliação. O data/val.jsonl (30 itens) é congelado: é
idêntico, item a item, ao do ciclo anterior, para que os dois modelos sejam
comparados de forma justa nas mesmas questões. Nenhum item novo entrou nele.
Um segundo conjunto, data/val_novos_v1.jsonl (30 itens, 3 por habilidade
nova, 1º ao 4º ano), fica fora do treino para medir a cobertura dos anos
novos.
Aumentação sintética (generate_synthetic.py, sem alterações neste
ciclo). Somamos 244 exemplos sintéticos (396 questões, pois alguns
exemplos agrupam 2 a 5 questões numa única lista "questoes": [...]). Eles
cobrem adição, subtração, multiplicação, divisão, porcentagem, potenciação e
frações (habilidades H07–H09). A resposta é sempre calculada em Python antes
de a questão ser montada, nunca por um LLM. Cada item reaproveita uma tupla
real (ano, habilidade, descrição, dificuldade) do banco. O bloco sintético é
gerado pelo código-fonte, não pelo banco, e por isso é idêntico byte a byte
entre ciclos.
Dataset de treino final: 963 exemplos (719 reais + 244 sintéticos).
Formato do exemplo (chat SFT, inalterado neste ciclo):
system: instrução fixa com o papel do modelo e o schema JSON.user:"Gere {quantidade} questão(ões) de matemática. Ano: {ano}. Habilidade: {habilidade} — {descrição}. Dificuldade: {dificuldade}."assistant:{"questoes": [{"enunciado", "alternativas" (A–E), "resolucao_passo_a_passo", "resposta_correta", "difficulty"}, ...]}
O modelo escreve resolucao_passo_a_passo antes de resposta_correta:
primeiro raciocina, depois se compromete com a letra.
Procedimento de treinamento
Usamos QLoRA (Dettmers et al., 2023) com Unsloth e o SFTTrainer do TRL. O modelo base fica congelado em 4-bit NF4 e apenas os adaptadores LoRA são treinados, o que cabe numa GPU de notebook (RTX 3060 Laptop, 6 GB de VRAM). Mantivemos os hiperparâmetros deste ciclo; nada nos experimentos pedia mudança.
| Hiperparâmetro | Valor |
|---|---|
| Modelo base | unsloth/Qwen3-1.7B (unsloth/qwen3-1.7b-unsloth-bnb-4bit) |
max_seq_length |
1024 |
| LoRA rank / alpha | 16 / 32 |
| LoRA dropout | 0,0 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Batch efetivo | 2 × 8 de acumulação de gradiente = 16 |
| Learning rate | 2e-4, cosseno, warmup de 5% |
| Épocas | até 3, early stopping por eval_loss (patience 2) |
| Otimizador | adamw_8bit, weight decay 0,01 |
| Precisão | bf16 |
| Loss | apenas nos tokens do assistant (train on completions) |
| Seed | 42 |
eval_loss por época: 0,5772 → 0,5362 (melhor, época 2) → 0,5414.
train_loss final: 0,4424.
Versões: TRL 0.24.0, Transformers 5.5.0, PyTorch 2.10.0.
Isolando o efeito da atualização de software
Entre os ciclos, o ambiente Python original se perdeu e foi reconstruído com um stack mais novo (torch 2.10, transformers 5.5, unsloth 2026.9.4). Para separar o efeito dos dados novos do efeito das bibliotecas novas, treinamos um modelo de controle com os mesmos 517 exemplos do baseline anterior, no stack novo, com o mesmo código e os mesmos hiperparâmetros. O dataset é a única variável.
| Métrica (conjunto congelado, n=30, pareado por seed) | Controle (dados antigos) | Candidato (dados novos) |
|---|---|---|
eval_loss (melhor época) |
0,5583 | 0,5362 (−4,0%) |
| Menções à palavra "figura" | 10,0% | 10,0% (idêntico: efeito do stack, não dos dados) |
Avaliação
Avaliamos em dois conjuntos, com geração real pelo modelo GGUF no llama.cpp,
grammar GBNF, temperature=0.7, top_p=0.8 e seeds pareadas por amostra.
Conjunto congelado (2º, 5º e 9º ano; n=30; mesmo do ciclo anterior)
| Métrica estrutural | Resultado |
|---|---|
| JSON válido / wrapper válido / schema completo | 100% |
resposta_correta válida (A–E) |
100% |
| Cinco alternativas distintas | 100% |
difficulty válida e aderente ao pedido |
100% |
| Consistência resposta ↔ resolução¹ | 100% (7/30 amostras verificáveis) |
| Dependência de visual ausente² | 3,3% (1/30); não significativo vs. baseline (McNemar p = 1,000) |
| Viés de gabarito (letra mais frequente) | 36,7% (baseline: 46,7%) |
| Falhas irrecuperáveis após best-of-N | 0 |
¹ Consistência é uma checagem determinística
(schema_utils.check_consistency, sem LLM). Ela extrai da resolução uma
expressão do tipo "a op b = r" e a compara com a alternativa indicada em
resposta_correta. Só se aplica quando essa expressão é encontrada (7 de 30
amostras aqui). Questões com raciocínio verbal ou resposta textual ou
geométrica ficam de fora. No 9º ano, a cobertura cai para 2/19, ou seja,
a maior parte das gerações do 9º ano não pode ser verificada
automaticamente. Tratamos isso como risco de qualidade em aberto (ver
Limitações).
² depende_de_visual_ausente_pct substitui a antiga
mencoes_figura_pct como métrica de decisão. A métrica antiga contava
qualquer ocorrência de palavras como "figura", "imagem", "gráfico" ou
"desenho", o que gerava falsos positivos, como "Desenho" listado como hobby
numa alternativa, ou uma questão autocontida que apenas pede ao aluno para
construir um gráfico de barras. A métrica nova
(schema_utils.DEPENDENCIA_VISUAL_PATTERN) procura referências dêiticas,
como "observe a figura abaixo" ou "conforme a tabela". Nos nossos 14
casos de teste ela não errou, mas 14 é uma amostra pequena. O pipeline de produção
(test_model.generate_validated) agora também penaliza e regenera candidatos
que dependem de um visual ausente.
Holdout de cobertura (1º ao 4º ano; n=30; fora do treino)
Este é o resultado central do ciclo: o desempenho nos anos que só o lote
L2-2026-09 trouxe.
| Controle (não viu esses anos) | Candidato (treinado neles) | |
|---|---|---|
| Verificáveis por máquina | 23/30 | 30/30 |
| 1º ano | 9/9 | 9/9 |
| 2º ano | 7/9 | 9/9 |
| 3º ano | 4/9 | 9/9 |
| 4º ano | 3/3 | 3/3 |
| Estrutura / dificuldade / consistência | 100% | 100% |
| Dependência de visual ausente | 0% | 0% |
Teste de McNemar pareado (7 vitórias, 0 derrotas): p = 0,016. É o único resultado estatisticamente significativo do ciclo. A melhora no conjunto congelado (2º, 5º e 9º ano) está dentro do ruído (p = 0,625), o que era esperado: fora o 2º ano, o lote novo não trouxe questões desses anos.
| Métrica de linguagem | Resultado |
|---|---|
| Perplexidade (resposta de referência) | não medida neste ciclo (gate G9, informativo) |
| Velocidade (CPU de desktop, GGUF, 4 threads) | Resultado |
|---|---|
| Vazão de geração | 26,5 tokens/s (baseline: 26,3, mesmo pipeline com filtro de dependência visual) |
| Latência total média (inclui carga do modelo) | 17,9 s |
Gates de promoção (src/promover_checkpoint.py)
| Gate | Critério | Resultado |
|---|---|---|
| G1 Contrato estrutural | ≥ 99% e ≥ baseline em 7 flags | 100% em todas |
| G2 Falhas após best-of-N | = 0 | 0 |
| G3 Consistência | ≥ baseline − 5 pp | 100% → 100% |
| G4 Dependência visual | sem piora estatisticamente significativa | 0% → 3,3% (McNemar p = 1,000) |
| G5 Aderência à dificuldade | ≥ baseline − 5 pp | 100% → 100% |
| G6 Viés de gabarito | ≤ baseline + 5 pp | 46,7% → 36,7% |
| G7 Esquecimento (5º e 9º ano) | sem regressão estrutural | preservado |
| G8 Velocidade | ≥ baseline − 10% | 26,3 → 26,5 tok/s |
| G9 Perplexidade (informativo) | — | não medida |
Decisão: PROMOVIDO. Os 8 gates bloqueantes foram aprovados.
Como usar
Adaptadores LoRA com Unsloth
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="<este-repo>/lora",
max_seq_length=1024,
load_in_4bit=True,
)
FastLanguageModel.for_inference(model)
messages = [
{"role": "system", "content": "Você é um gerador de questões de matemática no padrão SAEB..."},
{"role": "user", "content": "Gere 1 questão(ões) de matemática. Ano: 3º ano. Habilidade: EF03MA01 — ... . Dificuldade: Fácil."},
]
inputs = tokenizer.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
enable_thinking=False, return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512, temperature=0.7, top_p=0.8)
print(tokenizer.decode(output[0, inputs.shape[1]:], skip_special_tokens=True))
GGUF com llama.cpp (offline / mobile)
llama-cli -m gguf/qwen3-1.7b.Q4_K_M.gguf --temp 0.7 --top-p 0.8 \
--grammar-file grammars/questao.gbnf \
-p "Gere 1 questão(ões) de matemática. Ano: 3º ano. Habilidade: EF03MA07 — ... . Dificuldade: Moderado."
Recomendamos usar sempre a grammar GBNF (grammars/questao.gbnf). Ela
força o contrato exato: wrapper {"questoes": [...]}, cinco alternativas
(A–E), resposta_correta em {A,B,C,D,E} e difficulty em
{EASY,MEDIUM,HARD}. Combine-a com generate_validated()
(src/test_model.py), que verifica estrutura, consistência e dependência
visual antes de entregar a questão, regenerando quando necessário.
Limitações
- A verificação matemática ainda é fraca no 9º ano (2/19 amostras
verificáveis neste ciclo). Observamos um item com resolução
matematicamente errada: calculava
√(376,8/√3) ≈ 10√3, quando o valor correto é cerca de 14,7, e afirmava que um triângulo equilátero inscrito num círculo de raio 10 cm ocupa 30% da área do círculo (o valor correto é cerca de 41%). O pipeline o aceitou com statusnao_verificavel, por não haver uma expressão "a op b = r" reconhecível. Este é o maior risco de qualidade em aberto. Já existia no modelo anterior, portanto não é uma regressão, mas não foi corrigido neste ciclo. Um verificador simbólico (ex.:sympy) para as habilidades algébricas e geométricas do 9º ano é o próximo passo natural. - A letra da resposta pode discordar da resolução nas amostras que o
verificador determinístico não cobre. Isso é esperado num modelo de 1,7B
sem chain-of-thought explícito. Best-of-N e correção determinística
(
generate_validated()) reduzem o problema, mas não o eliminam. - Os dados reais estão concentrados. O lote
L2-2026-09representa 61% do corpus real, mas cobre só 10 habilidades (cerca de 48 itens cada), contra cerca de 5,5 itens em cada uma das 55 habilidades anteriores. O modelo pode se especializar demais nessas 10 habilidades; monitoramos isso comdata/val_novos_v1.jsonl. - O lote inverteu a distribuição de dificuldade. Questões fáceis subiram para 52% do corpus real (eram 27%) e difíceis caíram para 15% (eram 35%). A aderência à dificuldade se manteve em 100% nos dois conjuntos, mas a distribuição de treino ficou desbalanceada e precisa de atenção nos próximos ciclos.
- A alternativa E das questões reais é sempre um distrator fixo ("Nenhuma das alternativas anteriores"), porque o banco original só tem A–D. Apenas as questões sintéticas têm cinco distratores pedagogicamente distintos.
- O stack de treino não está pinado. O
requirements.txtnão fixa, por decisão de projeto, as versões de torch, unsloth e transformers, o que impediu reproduzir exatamente o ambiente do ciclo anterior. Recomendamos congelar umrequirements.locka partir do stack validado aqui (torch 2.10.0, transformers 5.5.0, unsloth 2026.9.4, trl 0.24.0). - Ainda não testado em hardware mobile real. Todas as medidas de velocidade vêm de CPU de desktop com 4 threads; latência e memória no dispositivo são desconhecidas.
- Geração em lote (
quantidade > 1questões por chamada) não foi avaliada neste ciclo, embora o contrato e os dados sintéticos a suportem.
Uso pretendido e considerações éticas
O modelo foi feito para apoiar o professor, não para substituí-lo. As questões geradas devem ser revisadas por um educador antes de chegar aos alunos, sobretudo no 9º ano, onde a verificação automática é limitada. Ele não foi projetado para avaliações de alto impacto e não gera questões que dependam de imagens.
Proveniência dos dados
As questões vêm de um banco interno de itens no padrão SAEB usado pelo
projeto (553 itens originais) e de um lote de 486 questões autorais escritas
por colaboradores conforme TreinamentoDados/GUIA_ENTREGA_ALUNOS.md. O lote
foi auditado e incorporado no ciclo L2-2026-09
(src/validar_entrega.py, src/importar_entregas.py). O código de
treinamento, extração e avaliação está disponível no repositório GitHub
indicado acima.
Citação
Se você usar este modelo, seu pipeline de dados ou seu protocolo de avaliação, cite:
@misc{sarmanho2026qwen3saeb,
author = {Sarmanho, Elton},
title = {{Qwen3-1.7B SAEB Math Question Generator}: On-Device
Generation of Structured Multiple-Choice Mathematics
Questions for Brazilian Basic Education},
year = {2026},
month = sep,
howpublished = {\url{https://github.com/eltonsarmanho/TreinamentoModeloQuestoes}},
note = {QLoRA fine-tune of Qwen3-1.7B, data cycle L2-2026-09,
GGUF Q4\_K\_M release}
}
Cite também os métodos de base:
@inproceedings{dettmers2023qlora,
author = {Dettmers, Tim and Pagnoni, Artidoro and Holtzman, Ari and Zettlemoyer, Luke},
title = {{QLoRA}: Efficient Finetuning of Quantized {LLMs}},
booktitle = {Advances in Neural Information Processing Systems},
volume = {36},
year = {2023}
}
@misc{qwen3technicalreport,
author = {{Qwen Team}},
title = {{Qwen3} Technical Report},
year = {2025},
eprint = {2505.09388},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}
- Downloads last month
- 169
4-bit