Download benchmarking.md from Wiself/Voice: direct link, hf CLI and curl.
- Browser
- Download file 5.69 kB
-
https://huggingface.co/Wiself/Voice/resolve/main/benchmarking.md
- Command line
-
hf download hf://Wiself/Voice/benchmarking.md
-
curl -L -o benchmarking.md https://huggingface.co/Wiself/Voice/resolve/main/benchmarking.md
Benchmarking methodology
Capability regression for base vs. voiced models. Standalone harness in
benchmarking/; only results and this methodology are published.
Engine
lm-evaluation-harness
(lm-eval==0.4.13, see benchmarking/requirements.txt) against a served
model over an OpenAI-compatible endpoint. No model weights are loaded by the
harness; inference stays in llama-server. Isolated venv, stdlib-only shim.
Adapter: benchmarking/shim.py
llama-server and lm-eval disagree on three shapes; the shim translates, nothing more:
GET /tokenizer_info— missing on llama-server. Served from the server's own/props(bos/eos strings, canonical chat template).POST /tokenize {"prompt"}→ upstream{"content"}→{"tokens"}.POST /detokenize— same translation, both directions./v1/completionslogprobs: server returns chat-stylecontent[{token, logprob, top_logprobs}]; harness expects legacy{token_logprobs[], top_logprobs[]}. Translated field-for-field.- Everything else is a transparent reverse proxy.
Tokenizer
There is no local tokenizer file and none is downloaded. All tokenization is performed by the served model itself via its own embedded tokenizer, which is exact by construction. The chat template is the server's canonical template for the served architecture (for Gemma 4 models: Google's canonical Gemma 4 chat template). No gated downloads, no vocabulary mismatch possible.
Tasks
| Task | Type | Items | Backend |
|---|---|---|---|
| tinyArc (logprob) | multiple-choice logprobs | 100 | ABANDONED, see below |
| tinyMMLU (logprob) | multiple-choice logprobs | 100 | ABANDONED, see below |
| arc_challenge_chat | generative letter answer | 100 (first of test) | local-chat-completions, server template |
| mmlu_generative | generative letter answer | 114 (2 × 57 subjects) | local-chat-completions, server template |
| tinyGSM8k | generation, exact match | 200 | local-chat-completions, server template |
Logprob scoring was abandoned: llama-server's /v1/completions does not
return prompt-token logprobs (echo unsupported), so multiple-choice
loglikelihood is unmeasurable on this endpoint. Generative letter-answer
scoring (the same style as OpenAI simple-evals MMLU) replaces it on every
task. Temperature 0.0 everywhere; greedy, deterministic.
Results
Two files, temperature 0.0 throughout, same prompts, budgets, and extraction on both:
- File A:
TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf, unmodified. - File B:
lm_head-TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf— file A with itslm_headtensor replaced viavoice cast, Q8_0. Nothing else changed.
| Task | n | A correct | B correct | A acc | B acc |
|---|---|---|---|---|---|
| ARC-Challenge (generative) | 100 | 94 | 95 | 0.940 | 0.950 |
| MMLU (generative, 2 × 57 subjects) | 114 | 101 | 103 | 0.886 | 0.904 |
| GSM8K | 100 | 76 | 92 | 0.760 | 0.920 |
Wrong/unanswered split — A: ARC 6/0, MMLU 5/8, GSM8K 1/23. B: ARC 5/0, MMLU 7/4, GSM8K 2/6. Unanswered = empty model content after thought exhausted the token budget (counted separately, never folded into wrong). GSM8K sample files contain each doc twice (harness duplication); scores dedup by doc_id, and split rescue pairs keep the decided verdict.
Run labeling
Each run is labeled file-a or file-b in the model name and output path.
A run record consists of the harness results_*.json plus samples_*.jsonl
(every item) plus the shim's thinking log (every prompt, answer, and thinking
trace). Harness stock filters cannot read thinking-model output
formats (echoed answer prefixes, thought-first responses), so quoted scores
come from benchmarking/rescore.py, which extracts the final answer letter
(answer is X preferred, else last bare A-D) or number (#### N, else last
integer). Unanswered items (empty content after thought exhausted the token
budget) are reported separately, never silently folded into wrong.
Three thinking-model adaptations, all documented here so results stay
comparable: (1) system messages are folded into the user turn (Gemma
supports only user/model roles); (2) generation stop-sequences are
["</s>"] — task-default stops on "\n" decapitate thought mid-stream;
(3) token budgets 1024 (ARC/GSM8K) / 2048 (MMLU) to leave room for thought.
Reproduce
uv venv benchmarking/.venv
uv pip install --python benchmarking/.venv/bin/python -r benchmarking/requirements.txt
python3 benchmarking/shim.py 8081
benchmarking/.venv/bin/python -m lm_eval run --model local-chat-completions
--model_args base_url=http://127.0.0.1:8081/v1/chat/completions,model=