Voice / benchmarking.md
Wiself's picture
small tiny fixes
271bd15
|
Raw History Blame Contribute Delete
5.69 kB

Benchmarking methodology

Capability regression for base vs. voiced models. Standalone harness in benchmarking/; only results and this methodology are published.

Engine

lm-evaluation-harness (lm-eval==0.4.13, see benchmarking/requirements.txt) against a served model over an OpenAI-compatible endpoint. No model weights are loaded by the harness; inference stays in llama-server. Isolated venv, stdlib-only shim.

Adapter: benchmarking/shim.py

llama-server and lm-eval disagree on three shapes; the shim translates, nothing more:

  • GET /tokenizer_info — missing on llama-server. Served from the server's own /props (bos/eos strings, canonical chat template).
  • POST /tokenize {"prompt"} → upstream {"content"} → {"tokens"}.
  • POST /detokenize — same translation, both directions.
  • /v1/completions logprobs: server returns chat-style content[{token, logprob, top_logprobs}]; harness expects legacy {token_logprobs[], top_logprobs[]}. Translated field-for-field.
  • Everything else is a transparent reverse proxy.

Tokenizer

There is no local tokenizer file and none is downloaded. All tokenization is performed by the served model itself via its own embedded tokenizer, which is exact by construction. The chat template is the server's canonical template for the served architecture (for Gemma 4 models: Google's canonical Gemma 4 chat template). No gated downloads, no vocabulary mismatch possible.

Tasks

Task Type Items Backend
tinyArc (logprob) multiple-choice logprobs 100 ABANDONED, see below
tinyMMLU (logprob) multiple-choice logprobs 100 ABANDONED, see below
arc_challenge_chat generative letter answer 100 (first of test) local-chat-completions, server template
mmlu_generative generative letter answer 114 (2 × 57 subjects) local-chat-completions, server template
tinyGSM8k generation, exact match 200 local-chat-completions, server template

Logprob scoring was abandoned: llama-server's /v1/completions does not return prompt-token logprobs (echo unsupported), so multiple-choice loglikelihood is unmeasurable on this endpoint. Generative letter-answer scoring (the same style as OpenAI simple-evals MMLU) replaces it on every task. Temperature 0.0 everywhere; greedy, deterministic.

Results

Two files, temperature 0.0 throughout, same prompts, budgets, and extraction on both:

  • File A: TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf, unmodified.
  • File B: lm_head-TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf — file A with its lm_head tensor replaced via voice cast, Q8_0. Nothing else changed.
Task n A correct B correct A acc B acc
ARC-Challenge (generative) 100 94 95 0.940 0.950
MMLU (generative, 2 × 57 subjects) 114 101 103 0.886 0.904
GSM8K 100 76 92 0.760 0.920

Wrong/unanswered split — A: ARC 6/0, MMLU 5/8, GSM8K 1/23. B: ARC 5/0, MMLU 7/4, GSM8K 2/6. Unanswered = empty model content after thought exhausted the token budget (counted separately, never folded into wrong). GSM8K sample files contain each doc twice (harness duplication); scores dedup by doc_id, and split rescue pairs keep the decided verdict.

Run labeling

Each run is labeled file-a or file-b in the model name and output path. A run record consists of the harness results_*.json plus samples_*.jsonl (every item) plus the shim's thinking log (every prompt, answer, and thinking trace). Harness stock filters cannot read thinking-model output formats (echoed answer prefixes, thought-first responses), so quoted scores come from benchmarking/rescore.py, which extracts the final answer letter (answer is X preferred, else last bare A-D) or number (#### N, else last integer). Unanswered items (empty content after thought exhausted the token budget) are reported separately, never silently folded into wrong.

Three thinking-model adaptations, all documented here so results stay comparable: (1) system messages are folded into the user turn (Gemma supports only user/model roles); (2) generation stop-sequences are ["</s>"] — task-default stops on "\n" decapitate thought mid-stream; (3) token budgets 1024 (ARC/GSM8K) / 2048 (MMLU) to leave room for thought.

Reproduce

uv venv benchmarking/.venv uv pip install --python benchmarking/.venv/bin/python -r benchmarking/requirements.txt python3 benchmarking/shim.py 8081 benchmarking/.venv/bin/python -m lm_eval run --model local-chat-completions
--model_args base_url=http://127.0.0.1:8081/v1/chat/completions,model=