Recipe
qwen3-6-35b-a3b-mq4r-rx-7900-xtx-24gb-hipfire-tp1
qwen3-6-35b-a3b-mq4r-rx-7900-xtx-24gb-hipfire-tp1Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- hipfire
- Engine version
- 0.3.0+72877258068c
- Accelerators
- 1
- Tensor parallel
- 1
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/Qwen/Qwen3.6-35B-A3B- Repository
- Qwen/Qwen3.6-35B-A3B
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
hipfire
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
hipfireserve--model~/.hipfire/models/qwen3.6-35b-a3b.mq4r--mtp~/.hipfire/models/qwen3.6-35b-a3b.mtp--mtp-ngram--ngram-match24--ngram-min48--ngram-max64--prompt-filebenchmarks/prompts/code_edit_rewrite_copy.txt--max-tokens256--max-seq4096--kvq8--temperature0--modebattery
| Flag | Value |
|---|---|
--model | ~/.hipfire/models/qwen3.6-35b-a3b.mq4r |
--mtp | ~/.hipfire/models/qwen3.6-35b-a3b.mtp |
--ngram-match | 24 |
--ngram-min | 48 |
--ngram-max | 64 |
--prompt-file | benchmarks/prompts/code_edit_rewrite_copy.txt |
--max-tokens | 256 |
--max-seq | 4096 |
--kv | q8 |
--temperature | 0 |
--mode | battery |
Environment
| Variable | Value |
|---|---|
DEVICE | 0 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 4,096 | 1,467.8 | 494.7 | 520.5 | observed | qwen3-6-35b-a3b-mq4r-rx-7900-xtx-24gb-hipfire-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- rx-7900-xtx-24gb
- id
- qwen3-6-35b-a3b-mq4r-rx-7900-xtx-24gb-hipfire-tp1
- model instance id
- qwen-qwen3-6-35b-a3b--mq4r
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen3-6-35b-a3b-mq4r-rx-7900-xtx-24gb-hipfire-tp1-sweep
- status
- candidate
capabilities
engine
- name
- hipfire
- version
- 0.3.0+72877258068c
serving
- tensor parallel
- 1
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
- serving.max concurrency · provenance · captured at
- 2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown
- serving.max context tokens · provenance · captured at
- 2026-09-01T10:20:39Z
serving.max context tokens · reason context-limit-not-evidenced
serving.max context tokens · state unknown
metadata
- localmaxxing · backend
- rocm
- localmaxxing · batch size
- 1
- localmaxxing · hardware label
- RX 7900 XTX
- localmaxxing · notes
- SINGLE-STREAM. Speculative decode active (MTP + n-gram). tau A 3.81 -> B 5.8 (2.81x), thresholds 24/48/64 defaults, ngram_mod 3 windows 192 drafts 192 accepted, parity md5 6eb9a69a (4/4 byte-identical output). engine: hipfire 0.3.0+72877258068c, backend rocm, gfx1100 quant: MQ4R = MagnumQuant 4-bit Redline (graded mixed-tier: MQ4 attn/router/shared + graded MQ4 routed experts, Redline retained-PM4 dispatch), 4.16 bpw effective (18700048128 bytes *8 / 35.95B params) model file: qwen3.6-35b-a3b.mq4r sha256 4685c140c46b1a6f prompt: code_edit_rewrite_copy.txt md5 80b910784c456bbb5532eb0493a497b9, 764 prompt tokens (templated) method: median of 4 runs (ABBA A B B A), batch1 single-stream, fresh process, temperature 0 (greedy) max_tokens 256 ctx 4096 KV q8 evidence: sealed case a3b-mq4r-ngram-edit-abba-20260810-gfx1100-rx7900xtx, corpus manifest manifest-sha12-8f3c2e1a4b9d not measured: peak VRAM (not instrumented this campaign)
- localmaxxing · observed command
- DEVICE=0 hipfire serve --model ~/.hipfire/models/qwen3.6-35b-a3b.mq4r --mtp ~/.hipfire/models/qwen3.6-35b-a3b.mtp --mtp-ngram --ngram-match 24 --ngram-min 48 --ngram-max 64 --prompt-file benchmarks/prompts/code_edit_rewrite_copy.txt --max-tokens 256 --max-seq 4096 --kv q8 --temperature 0 --mode battery
- localmaxxing · run id
- cmsnp3cmd00nyo001yvkgm3tq
- localmaxxing · tokenized · arguments
- hipfire, serve, --model, ~/.hipfire/models/qwen3.6-35b-a3b.mq4r, --mtp, ~/.hipfire/models/qwen3.6-35b-a3b.mtp, --mtp-ngram, --ngram-match, 24, --ngram-min, 48, --ngram-max, 64, --prompt-file, benchmarks/prompts/code_edit_rewrite_copy.txt, --max-tokens, 256, --max-seq, 4096, --kv, q8, --temperature, 0, --mode, battery
- localmaxxing · tokenized · environment · DEVICE
- 0
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmsnp3cmd00nyo001yvkgm3tq ↗ |