Benchmarks

Proof on silicon, not slides

FastFlowLM is tuned on real Ryzen™ AI hardware with synthetic and application-level workloads. Expect steady 40–80 tok/s on 7B models at < 10 W, plus deterministic latency for agentic chains.

  • Full-stack telemetry

    Counters for NPU, CPU, and memory let you see exactly where cycles go.

  • Scenario-driven suites

    Instruction tuning, RAG, chat, and multimodal tests mirror real workloads.

Llama3.2 3B @ 4-bit

72 tok/s

Ryzen™ AI 9 HX 370 · 8 ms median latency

Gemma 3 4B Vision

18 fps

Vision + text pipeline on XDNA2 with shared memory.

Power draw

9.6 W

Full assistant stack vs ~45 W GPU baseline.