Download results.md from Wiself/Voice: direct link, hf CLI and curl.
- Browser
- Download file 3.47 kB
-
https://huggingface.co/Wiself/Voice/resolve/main/results.md
- Command line
-
hf download hf://Wiself/Voice/results.md
-
curl -L -o results.md https://huggingface.co/Wiself/Voice/resolve/main/results.md
Results: two files, one tensor apart
Date: 2026-09-04. Two runs of the same benchmarks against two model files
at temperature 0.0. Method: benchmarking.md
(harness, settings, adaptations, extraction rules — everything needed
to reproduce).
Setup
- File A:
TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf(Gemma 4 26B MoE, IQ4), served via llama-server, unmodified. - File B:
lm_head-TheDrummer_Orion-26B-A4B-v1-IQ4_NL.gguf— file A with itslm_headtensor replaced viavoice cast, Q8_0. Nothing else changed. - Both files: temperature 0.0, seed 1234, identical prompts, identical token
budgets (ARC 1024, MMLU 2048, GSM8K 1024), identical answer extraction.
Unanswered items got one rescue pass at 4096 (MMLU) / 2048 (GSM8K),
applied to both files symmetrically. Rescue tie-break: on split pairs
(two verdicts for one doc) the decided verdict beats
unanswered.
Scores
| Task | n | A correct | B correct | A acc | B acc |
|---|---|---|---|---|---|
| ARC-Challenge (generative) | 100 | 94 | 95 | 0.940 | 0.950 |
| MMLU (generative, 2 × 57 subjects) | 114 | 101 | 103 | 0.886 | 0.904 |
| GSM8K | 100 | 76 | 92 | 0.760 | 0.920 |
Observations
- ARC (+1) and MMLU (+2) differences fall within slice noise (±5 points at n≈100). GSM8K (+16) exceeds it; see point 2.
- The GSM8K difference sits in unanswered counts. File A left 23 items silent (empty content after thought exhausted the budget); file B left 6. On items both files answered, accuracy is 0.987 (A) vs 0.979 (B). The traces hold whatever else there is to see.
- Thinking lengths roughly match. No pairing script was archived, so no median is claimed; per-file thinking lengths sit in the same range.
- Error overlap is minimal. One GSM8K item was wrong under both files; remaining errors are different items per file.
What this does not show
- Small slices: ~100 items per task. Differences under ~5 points are noise.
- Generative scoring is not comparable to published logprob leaderboards; compare only file A vs. file B within this table.
- One model pair, one replaced tensor, greedy decoding only. Other pairs
may differ — that is what the harness in
benchmarking/is for. - Writing voice is measured separately below — same two files, same neutrality, no judge model.
Evidence
Per file, archived under benchmarking/runs/file-a/ and
benchmarking/runs/file-b/: harness results_*.json and
samples_*.jsonl for every run, rescue files for the second passes, and
the thinking log (every prompt, answer, and reasoning trace). No personal
information in any of it — server file paths are redacted to run labels
at the log source.
Part II — writing voice: 32 prompts, two files
Method: creative-bench.md. 32 EQ-Bench v3 prompts,
one generation per prompt per file at temperature 0.7, 6000-token budget.
Measured, not judged: shared trigram vocabulary, cliche density, texture
stats. Blind A/B pairs archived for readers.
| Measure | File A | File B |
|---|---|---|
| Shared trigram vocab, mean (max) | 0.013 (0.027) across all 32 | same pair |
| Cliche hits per 100 words | 0.141 | 0.075 |
| Avg words per piece | 1018.0 | 878.6 |
| Lexical diversity (TTR) | 0.4439 | 0.5054 |
| Avg sentence words | 9.3 | 8.8 |
Per-prompt exhibits: creative-bench/runs/file-a-creative.jsonl and
file-b-creative.jsonl (prompt, full text, full reasoning trace each).