Abschrift
Docs / Local notes benchmark

Local meeting notes, benchmarked

This page is generated from the same benchmark results that generate the model catalog shipped inside the app — the numbers you read are the numbers the app acts on.

Abschrift writes meeting notes on your Mac by default: a built-in runtime plus an open model picked for your hardware, downloaded once. That promise is only honest if the local model is actually good. So we measure it — with a benchmark built around the one job that matters: turning a messy meeting transcript into notes you can act on. This edition covers 17 open models from 2B to 122B parameters, from an 8 GB MacBook Air to a 128 GB Mac Studio.

97.4
Frontier baseline (Claude Fable 5) — the bar
91.5
Best local score — Muse Glimmer 30B
17
Open models benchmarked, 2B–122B

What we measure

Six synthetic meetings (no real people or companies), written to be hard in the ways real meetings are hard: short stand-ups to two-hour strategy workshops, English and German, 4–23k tokens of transcript. Each comes with a scripted ground truth: every fact, decision, and action item that correct notes must carry, with owners and deadlines. Each also contains noise that good notes must drop, like small talk and venting. And each sets attribution traps: heated accusations whose substance belongs in the notes but whose finger-pointing does not.

Every model gets the app's actual notes prompt and identical sampling. Models up to 12B ran on the app's actual bundled runtime (a pinned llama-server build, Q4-class quantization) on a base M1, so speed is measured for real. Models too big for the measurement machine were generated through the OpenRouter API at provider precision and are marked “API” below. Their Apple-Silicon speeds are estimates. The methodology section at the end explains how.

A frontier model, Claude Fable 5, grades each output against the ground truth and the full transcript: fact coverage, false claims weighted by severity, action-item fidelity, section structure, attribution neutrality, noise rejection, and whether the Discussion section alone could power follow-up work. Two grading passes per output, averaged. The same rubric grades notes written by Claude Fable 5 itself. That's the bar, and it scores 97.4.

Does a bigger model write better notes?

4050607080901002B4B8B15B30B60B120BFrontier baseline (Claude Fable 5) — 97.4“excellent” bar — 85Qwen3.5 2BQwen3.5 2Bcomposite 48.6 · long meetings 50.62B params · 1.3 GB downloadquality measured on M1Qwen3.5 4BQwen3.5 4Bcomposite 77.3 · long meetings 65.04B params · 2.7 GB downloadquality measured on M1Gemma 4 E4BGemma 4 E4Bcomposite 76.3 · long meetings 75.04B params · 5.2 GB downloadquality measured on M1Granite 4.1 8BGranite 4.1 8Bcomposite 60.5 · long meetings 50.08B params · 5.3 GB downloadquality measured on M1Qwen3.5 9BQwen3.5 9Bcomposite 86.1 · long meetings 80.99B params · 5.7 GB downloadquality measured on M1Gemma 4 12BGemma 4 12Bcomposite 73.0 · long meetings 67.112B params · 7.0 GB downloadquality measured on M1Ministral 3 14BMinistral 3 14Bcomposite 84.5 · long meetings 83.814B params · 8.2 GB downloadquality graded via APIgpt-oss 20Bgpt-oss 20B (MoE)composite 72.2 · long meetings 78.821B params · 12.1 GB downloadquality graded via APIGemma 4 26B A4BGemma 4 26B A4B (MoE)composite 76.0 · long meetings 66.826B params · 14.4 GB downloadquality graded via APIQwen3.5 27BQwen3.5 27Bcomposite 89.3 · long meetings 90.127B params · 16.7 GB downloadquality graded via APIQwen3.8 27BQwen3.8 27Bcomposite 89.9 · long meetings 94.327B params · 16.5 GB downloadquality graded via APIMuse Glimmer 30BMuse Glimmer 30Bcomposite 91.5 · long meetings 86.630B params · 15.9 GB downloadquality graded via APIGemma 4 31BGemma 4 31Bcomposite 75.8 · long meetings 65.331B params · 17.6 GB downloadquality graded via APIQwen3.5 35B A3BQwen3.5 35B A3B (MoE)composite 86.6 · long meetings 85.835B params · 22.0 GB downloadquality graded via APIGLM-4.5 Air 106BGLM-4.5 Air 106B A12B (MoE)composite 88.5 · long meetings 81.2106B params · 73.0 GB downloadquality graded via APIgpt-oss 120Bgpt-oss 120B (MoE)composite 86.8 · long meetings 84.4117B params · 63.4 GB downloadquality graded via APIQwen3.5 122B A10BQwen3.5 122B A10B (MoE)composite 88.8 · long meetings 88.2122B params · 76.5 GB downloadquality graded via APIModel size (parameters, log scale)Notes quality (composite 0–100)
QwenGemmagpt-ossGLMMistralGraniteMeta
Notes quality vs model size, all 17 models. Hover any point for details. Log-scale x-axis; the dashed line is the frontier cloud baseline.

Results

ModelParamsCompositeLong meetings FactsCorrectnessAction items Quality viaVerdict
Frontier baseline (Claude Fable 5)97.498.698.298.397.3The bar
Muse Glimmer 30B30B91.586.692.9100.087.3APIExcellent
Qwen3.8 27B27B89.994.395.895.076.5APIExcellent
Qwen3.5 27B27B89.390.189.0100.081.5APIExcellent
Qwen3.5 122B A10B (MoE)122B (A10B)88.888.291.297.578.8APIExcellent
GLM-4.5 Air 106B A12B (MoE)106B (A12B)88.581.284.197.587.1APIExcellent
gpt-oss 120B (MoE)117B (A5B)86.884.484.193.383.7APIExcellent
Qwen3.5 35B A3B (MoE)35B (A3B)86.685.888.197.575.2APIExcellent
Qwen3.5 9B9B86.180.986.196.769.9on-deviceExcellent
Ministral 3 14B14B84.583.888.682.965.1APIGood
Qwen3.5 4B4B77.365.073.775.083.8on-deviceGood
Gemma 4 E4B4B76.375.056.785.091.1on-deviceGood
Gemma 4 26B A4B (MoE)26B (A4B)76.066.855.282.995.0APIGood
Gemma 4 31B31B75.865.362.282.576.7APIGood
Gemma 4 12B12B73.067.154.075.490.3on-deviceAcceptable
gpt-oss 20B (MoE)21B (A4B)72.278.864.584.652.1APIAcceptable
Granite 4.1 8B8B60.550.039.967.555.9on-deviceNot recommended
Qwen3.5 2B2B48.650.642.710.852.4on-deviceNot recommended

Scores are 0–100 composites (weights: facts 30%, correctness 20%, action items 15%, discussion quality 15%, attribution 10%, structure 5%, noise 5%). Long meetings is the composite over the 90–110-minute transcripts only — where small models fall behind first. (A…B) marks mixture-of-experts models: total (active) parameters.

Quality × speed: the map that picks your model

40506070809010010204080150300fast and excellentexcellent, needs patienceQwen3.5 2BQwen3.5 2Bcomposite 48.6~108.3 tok/s on M4 Prospeed measured on M1, scaledQwen3.5 4BQwen3.5 4Bcomposite 77.3~50.1 tok/s on M4 Prospeed measured on M1, scaledQwen3.5 9BQwen3.5 9Bcomposite 86.1~33.6 tok/s on M4 Prospeed measured on M1, scaledGranite 4.1 8BGranite 4.1 8Bcomposite 60.5~26.5 tok/s on M4 Prospeed measured on M1, scaledGemma 4 E4BGemma 4 E4Bcomposite 76.3~57.6 tok/s on M4 Prospeed measured on M1, scaledGemma 4 12BGemma 4 12Bcomposite 73.0~24.3 tok/s on M4 Prospeed measured on M1, scaledMinistral 3 14BMinistral 3 14Bcomposite 84.5~23.6 tok/s on M4 Prospeed estimated (±40%)gpt-oss 20Bgpt-oss 20B (MoE)composite 72.2~61.2 tok/s on M4 Prospeed community-measured anchorQwen3.5 27BQwen3.5 27Bcomposite 89.3~11.6 tok/s on M4 Prospeed estimated (±40%)Qwen3.8 27BQwen3.8 27Bcomposite 89.9~11.8 tok/s on M4 Prospeed estimated (±40%)Gemma 4 26B A4BGemma 4 26B A4B (MoE)composite 76.0~54.7 tok/s on M4 Prospeed estimated (±40%)Gemma 4 31BGemma 4 31Bcomposite 75.8~11.0 tok/s on M4 Prospeed estimated (±40%)Qwen3.5 35B A3BQwen3.5 35B A3B (MoE)composite 86.6~64.4 tok/s on M4 Prospeed estimated (±40%)GLM-4.5 Air 106BGLM-4.5 Air 106B A12B (MoE)composite 88.5~14.7 tok/s on M4 Prospeed estimated (±40%)gpt-oss 120Bgpt-oss 120B (MoE)composite 86.8~42.9 tok/s on M4 Prospeed community-measured anchorQwen3.5 122B A10BQwen3.5 122B A10B (MoE)composite 88.8~19.4 tok/s on M4 Prospeed estimated (±40%)Muse Glimmer 30BMuse Glimmer 30Bcomposite 91.5~12.2 tok/s on M4 Prospeed estimated (±40%)Generation speed on an M4 Pro (tokens/s, log scale)Notes quality (composite 0–100)
QwenGemmagpt-ossGLMMistralGraniteMeta
The ideal model sits top-right: frontier-approaching notes, instant generation. Speeds shown for an M4 Pro. Measured models scale from the M1 measurement by each chip's llama.cpp throughput. API-graded models are estimated (±25% dense / ±40% MoE).

What your Mac gets

The app maps unified memory to a default model. It only asks you to choose when the benchmark shows a genuine quality/speed trade-off on your tier:

60708090100frontier 97.48 GB+Qwen3.5 4B · 77.316 GB+Qwen3.5 9B · 86.132 GB+Qwen3.8 27B · 89.936 GB+Qwen3.8 27B · 89.996 GB+Qwen3.8 27B · 89.9128 GB+Qwen3.8 27B · 89.9
Default model per unified-memory tier, colored by model family; bar length = notes quality. The curve flattens: past the 27B class, more RAM buys smaller quality gains.
Unified memoryDefault modelDownload Compositeest. tok/s M1 / M2 / M4 / M4 Pro / M4 Max Alternatives
8 GB+Qwen3.5 4B2.74 GB77.314.0 / 21.6 / — / — / —
16 GB+Qwen3.5 9B5.68 GB86.19.4 / 14.5 / 16.0 / — / —Qwen3.5 4B (speed), Gemma 4 E4B (speed)
32 GB+Qwen3.8 27B16.46 GB89.9— / — / 5.6 / 11.8 / —Qwen3.5 9B (speed), Ministral 3 14B (speed)
36 GB+Qwen3.8 27B16.46 GB89.9— / — / — / 11.8 / 19.3Qwen3.5 35B A3B (MoE) (speed), Qwen3.5 9B (speed)
96 GB+Qwen3.8 27B16.46 GB89.9— / — / — / — / 19.3gpt-oss 120B (MoE) (speed), Qwen3.5 35B A3B (MoE) (speed)
128 GB+Qwen3.8 27B16.46 GB89.9— / — / — / — / 19.3Qwen3.5 122B A10B (MoE) (speed), gpt-oss 120B (MoE) (speed)

Blank cells mean the chip never shipped with that much RAM. Models above ~50 GB download in multiple parts. The app handles that automatically, each part checksum-pinned.

Why isn't Muse Glimmer 30B (composite 91.5) offered in the app? Its on-device runtime path has not been validated on real hardware yet and its mandatory reasoning makes notes take ~2.4x as long as the tier default's — it stays benchmark-only until that changes.

How long does a meeting take?

End-to-end notes generation for a ~45-minute meeting:

ModelNotes for a 45-min meeting
Muse Glimmer 30B≈23.2 min (estimated incl. mandatory reasoning, slowest capable Mac: M4)
Qwen3.8 27B≈12.2 min (estimated, slowest capable Mac: M4)
Qwen3.5 27B≈10.2 min (estimated, slowest capable Mac: M4)
Qwen3.5 122B A10B (MoE)≈1.6 min (estimated, slowest capable Mac: M4 Max)
GLM-4.5 Air 106B A12B (MoE)≈1.8 min (estimated, slowest capable Mac: M4 Max)
gpt-oss 120B (MoE)≈1.4 min (estimated incl. mandatory reasoning, slowest capable Mac: M2 Max)
Qwen3.5 35B A3B (MoE)≈0.7 min (estimated, slowest capable Mac: M3 Max)
Qwen3.5 9B5.2 min (measured, M1)
Ministral 3 14B≈8.0 min (estimated, slowest capable Mac: M3)
Qwen3.5 4B3.2 min (measured, M1)
Gemma 4 E4B2.9 min (measured, M1)
Gemma 4 26B A4B (MoE)≈1.5 min (estimated, slowest capable Mac: M3)
Gemma 4 31B≈6.2 min (estimated, slowest capable Mac: M4)
Gemma 4 12B7.0 min (measured, M1)
gpt-oss 20B (MoE)≈10.2 min (estimated incl. mandatory reasoning, slowest capable Mac: M3)
Granite 4.1 8B5.0 min (measured, M1)
Qwen3.5 2B1.3 min (measured, M1)

Methodology & honesty