Qwen3-8B-AWQ · TP=1
Profiles, resources and evidence for this configuration, followed by quality and token observations.
- Run reference
- Qwen3-8B-AWQ-vllm-xpu-tp1-baseline
- Engine
- vLLM-XPU live
- Topology / device
- TP=1 · Intel Arc Pro B70 (32GB)
- Quantization
- AWQ
- Execution
- Exclusive use
Qwen3-8B-AWQ · TP=110CHAPTERS
Measurement profiles · workload
Measurement profiles
Four distinct workloads with their observed values and conditions.
Decode · output length
Decode · output length
Observed raw values · medians separately in the table.
| Point | ISL | OSL | Raw value | TTFT |
|---|---|---|---|---|
| OSL 128 | 128 | 128 | 100.04 tok/s | 70.21 ms |
Raw request identity and complete workload
Statistics use only explicit inputs. Missing statistics remain not measured.
| Point | ISL | OSL | Context | C | Decode | Aggregate decode | Per-request decode | TTFT |
|---|---|---|---|---|---|---|---|---|
| Qwen3-8B-AWQ-vllm-xpu-tp1-baseline | 128 | 128 | 16,384 | 1 | 100.04 tok/s | 100.04 tok/s | 100.04 tok/s | 70.21 ms |
TTFT
Not measured for this run.
Context length
Not measured for this run.
Concurrency
Concurrency · aggregate throughput
Observed raw values · medians separately in the table.Per-request values are listed in raw details.
Categorical axis · equal spacing between measurement points.
| Point | C | OSL | Raw value | TTFT |
|---|---|---|---|---|
| C=1 | 1 | 128 | 100.04 tok/s | 70.21 ms |
| C=2 | 2 | 128 | 198.02 tok/s | 249.76 ms |
| C=4 | 4 | 128 | 381.13 tok/s | 125.45 ms |
| C=8 | 8 | 128 | 731.91 tok/s | 935.34 ms |
| C=16 | 16 | 128 | 1,201.58 tok/s | 229.15 ms |
| C=32 | 32 | 128 | 1,826.59 tok/s | 352.09 ms |
| C=64 | 64 | 128 | 2,485.19 tok/s | 613.38 ms |
| C=128 | 128 | 128 | 2,922.95 tok/s | 1,125.27 ms |
| C=256 | 256 | 128 | 3,242.01 tok/s | 2,209.05 ms |
| C=512 | 512 | 128 | 3,226.24 tok/s | 7,202.97 ms |
| C=1,024 | 1,024 | 128 | 3,203.07 tok/s | 17,202.16 ms |
Raw request identity and complete workload
Statistics use only explicit inputs. Missing statistics remain not measured.
| Point | ISL | OSL | Context | C | Aggregate decode | Per-request decode | TTFT |
|---|---|---|---|---|---|---|---|
| C=1 | 128 | 128 | 16,384 | 1 | 100.04 tok/s | 100.04 tok/s | 70.21 ms |
| C=2 | 128 | 128 | 16,384 | 2 | 198.02 tok/s | 100.21 tok/s | 249.76 ms |
| C=4 | 128 | 128 | 16,384 | 4 | 381.13 tok/s | 97.35 tok/s | 125.45 ms |
| C=8 | 128 | 128 | 16,384 | 8 | 731.91 tok/s | 92.18 tok/s | 935.34 ms |
| C=16 | 128 | 128 | 16,384 | 16 | 1,201.58 tok/s | 85.53 tok/s | 229.15 ms |
| C=32 | 128 | 128 | 16,384 | 32 | 1,826.59 tok/s | 67.3 tok/s | 352.09 ms |
| C=64 | 128 | 128 | 16,384 | 64 | 2,485.19 tok/s | 48.12 tok/s | 613.38 ms |
| C=128 | 128 | 128 | 16,384 | 128 | 2,922.95 tok/s | 29.26 tok/s | 1,125.27 ms |
| C=256 | 128 | 128 | 16,384 | 256 | 3,242.01 tok/s | 16.93 tok/s | 2,209.05 ms |
| C=512 | 128 | 128 | 16,384 | 512 | 3,226.24 tok/s | 15.16 tok/s | 7,202.97 ms |
| C=1024 | 128 | 128 | 16,384 | 1,024 | 3,203.07 tok/s | 14.15 tok/s | 17,202.16 ms |
Benchmark module · resources
Memory, prefix cache and energy
File size, observed GPU memory and reserved KV cache remain separate. Energy needs a device, phase, interval and source.
Memory by phase and device
Checkpoint on disk: Not measured
| Device | Idle | Post-load | KV reserved | Peak | Source |
|---|---|---|---|---|---|
| Memory | Not measured | ||||
Prefix-Cache · Cold / Warm
A comparison needs a matching cold/warm pair for the same prefix.
- Cold
- Not measured
- Warm
- Not measured
Energy by device and interval
- Device
- Not measured
- Phase
- Not measured
- Energy
- Not measured
- Power
- Not measured
- Duration
- Not measured
- Source
- Not measured
Benchmark module · evidence
Measurement artifact and evidence
A validation report belongs to an exact artifact digest. Raw artifacts, publication derivatives and reproduction details have different roles.
Artifact and validation report
- SHA-256
- live-capture
Raw artifact: Unavailable
Validation report: unvalidated_live_capture
Reproduction and identity
- Run-ID
- Qwen3-8B-AWQ-vllm-xpu-tp1-baseline
- Model revision
- Unavailable
- Measured at
- 2026-09-08T15:00:43.270497+00:00
- Source
- Unavailable
Reproduction details: Unavailable.
Model quality · comparison context
Model quality and quantization
Quality, throughput and memory apply to the supplied model version, quantization and evaluation method.
Select evaluation
Evaluation pending. No evaluation data is available for this run.
- Quality
- Not measured
- Throughput
- Not measured
- Memory
- Not measured
Comparison · three views
Quality × speed · three views
Three views of the same quality and throughput values. Absolute axes retain their units; normalization is only a presentation aid.
Evaluation pending. No evaluation data is available for this run.
Token observations
Speculative decoding
Requested behavior and effective behavior are separate facts. A counter does not replace a token trace.
Token observations unavailable for this run.
Requested vs. effective
- Requested
- Not supplied
- Effective
- Not supplied
- Proposed
- Not measured
- Accepted
- Not measured
- Rejected
- Not measured
- Dropped
- Not measured
Correctness cases
No correctness cases supplied.
Token observations
Acceptance map
Accepted, rejected and unused proposals after a rejection remain distinct.
Token observations unavailable for this run.
Token observations
Token latency
The sequence shows individual token events with observed latency. Throughput aggregates cannot produce a token timeline.
Token observations unavailable for this run.
Round × token
Token heatmap
Independent latency observations by round and token. Rejected and unevaluated cells remain distinct.
Not measured for this run