Models

One CLI, every Ryzen-ready model

Pull curated FastFlowLM recipes. The runtime streams tokens via HTTP, WebSocket, or the Ollama-compatible API, so existing apps work without rewrites.

Flagship reasoning

Llama 3.2 · DeepSeek · Qwen 3

Optimized kernels for 70B down to 1B, with automatic quantization and smart context reuse.

Vision & speech

Gemma 3 VLM · Whisper · Gemma Audio

VLM and audio pipelines render directly on the NPU, enabling private multimodal assistants.

Edge fine-tuning

FLM MoE + Embedding suites

Use built-in adapters, LoRA checkpoints, and embedding endpoints for retrieval workflows.