vLLM-XPU Measurements & evidence

Qwen2.5-Coder-7B-Instruct · TP=1

Profiles, resources and evidence for this configuration, followed by quality and token observations.

Run reference
Qwen2.5-Coder-7B-Instruct-vllm-xpu-tp1-baseline
Engine
vLLM-XPU live
Topology / device
TP=1 · Intel Arc Pro B70 (32GB)
Quantization
BF16
Execution
Exclusive use
Model → token sequence → evidence. Schematic illustration; no measurements.

Measurement profiles · workload

Measurement profiles

Four distinct workloads with their observed values and conditions.

Decode · output length

Decode · output length

Observed raw values · medians separately in the table.

ISL
128
OSL
128
C
1
Cache
not supplied
PointISLOSLRaw valueTTFT
OSL 12812812838.56 tok/s64.35 ms
Raw request identity and complete workload

Statistics use only explicit inputs. Missing statistics remain not measured.

PointISLOSLContextCDecodeAggregate decodePer-request decodeTTFT
Qwen2.5-Coder-7B-Instruct-vllm-xpu-tp1-baseline12812816,384138.56 tok/s38.56 tok/s38.56 tok/s64.35 ms

Benchmark module · resources

Memory, prefix cache and energy

File size, observed GPU memory and reserved KV cache remain separate. Energy needs a device, phase, interval and source.

Memory by phase and device

Checkpoint on disk: Not measured

DeviceIdlePost-loadKV reservedPeakSource
MemoryNot measured

Prefix-Cache · Cold / Warm

A comparison needs a matching cold/warm pair for the same prefix.

Cold
Not measured
Warm
Not measured

Energy by device and interval

Device
Not measured
Phase
Not measured
Energy
Not measured
Power
Not measured
Duration
Not measured
Source
Not measured

Benchmark module · evidence

Measurement artifact and evidence

A validation report belongs to an exact artifact digest. Raw artifacts, publication derivatives and reproduction details have different roles.

01 · PlanRun binding
02 · ManifestRun binding
03 · Raw artifactRun binding
04 · Validation reportRun binding
05 · DerivativeRun binding

Artifact and validation report

SHA-256
live-capture

Raw artifact: Unavailable

Validation report: unvalidated_live_capture

Reproduction and identity

Run-ID
Qwen2.5-Coder-7B-Instruct-vllm-xpu-tp1-baseline
Model revision
Unavailable
Measured at
2026-09-08T14:57:51.478905+00:00
Source
Unavailable

Reproduction details: Unavailable.

Model quality · comparison context

Model quality and quantization

Quality, throughput and memory apply to the supplied model version, quantization and evaluation method.

Model: Qwen2.5-Coder-7B-InstructTP: TP=1Engine: vLLM-XPU live

Select evaluation

Evaluation pending. No evaluation data is available for this run.

Quality
Not measured
Throughput
Not measured
Memory
Not measured

Comparison · three views

Quality × speed · three views

Three views of the same quality and throughput values. Absolute axes retain their units; normalization is only a presentation aid.

Evaluation pending. No evaluation data is available for this run.

01 · ABSOLUTE VALUES

Dual-axis lines

Quality (%) and throughput (tok/s) with separate, labeled axes.

Quality and throughput series unavailable.

02 · ABSOLUTE VALUES

Throughput bars + quality line

Quality (%) and throughput (tok/s) with separate, labeled axes.

Throughput bars and quality line unavailable.

03 · INDEX

Normalized comparison

Each series maximum = 100%. Relative shapes, without a shared measurement unit.

Normalization unavailable without observations.

Token observations

Speculative decoding

Requested behavior and effective behavior are separate facts. A counter does not replace a token trace.

Token observations unavailable for this run.

Requested vs. effective

Requested
Not supplied
Effective
Not supplied
Proposed
Not measured
Accepted
Not measured
Rejected
Not measured
Dropped
Not measured

Correctness cases

No correctness cases supplied.

Token observations

Acceptance map

Accepted, rejected and unused proposals after a rejection remain distinct.

Token observations unavailable for this run.

AcceptedRejectedDropped

Token observations

Token latency

The sequence shows individual token events with observed latency. Throughput aggregates cannot produce a token timeline.

Token observations unavailable for this run.

Round × token

Token heatmap

Independent latency observations by round and token. Rejected and unevaluated cells remain distinct.

Not measured for this run