LLMs on an Intel NPU
How I got Qwen3 8B running on my laptop's Intel AI Boost NPU with OpenVINO Model Server (or NoLlama), then hooked it into OpenWork and Claude Code, so small tasks stop burning Claude tokens and the GPU stays free.
Fuckass models didn't work, had to dig and find models that would work.
Abstract
I have an OMEN MAX Gaming Laptop 16 and recently I have been working on a lot of side projects in my free time. I've also been using Claude a lot since I got this laptop to become Peter Thiel's most feared 10x engineer.
Unfortunately, Claude has the habit of spawning more workflows, and that just absolutely DRAINS my token usage and limits.
Eventually I wanted to use a locally hosted model for smaller tasks and found out my laptop had an Intel AI Boost NPU that I could use instead of my GPU. This lets me run smaller LLM workloads without tying up the GPU while I am doing other work.
Then I realized there was very little documentation explaining how to actually use this POS hardware chip for an LLM. After a lot of experimenting, I managed to get Qwen running on the NPU through OpenVINO.
This page documents the setup I ended up using for OpenVINO Model Server, NoLlama, OpenWork, and Claude Code.
Backend Model Server
The frontend AI application and the model server are separate pieces.
Frontend
|
| OpenAI / Anthropic compatible API
v
Model Server
|
v
OpenVINO
|
v
Intel NPUFor the backend, I have mainly used OpenVINO Model Server and NoLlama.
OpenVINO Model Server
Prerequisites
- Python 3.12.X - Optional
OpenVINO Model Server has Windows packages with Python enabled and Python disabled. Some additional features and model workflows may require the Python-enabled package.
- Intel NPU Drivers
Download and install the latest Intel NPU drivers from Intel:
Intel NPU Driver for Windows
Installation
Download OpenVINO Model Server from the official repository:
https://github.com/openvinotoolkit/model_serverExtract OVMS somewhere on your computer.
For example:
C:\Users\YOUR_USERNAME\Downloads\ovmsOpen Command Prompt inside the OVMS directory and initialize the OpenVINO environment:
setupvars.bat
Useful OVMS Command Line Arguments
--rest_port 8000Starts the HTTP REST API on port 8000. The OpenAI-compatible API is exposed through this server.
--rest_bind_address 127.0.0.1Binds OVMS only to the local machine instead of exposing it to other devices on the network.
--rest_workers NControls the number of HTTP REST worker threads.
--port 9000Enables the gRPC server on port 9000.
--grpc_workers NControls the number of gRPC worker threads.
--pullDownloads the requested model before starting the server.
--source_model MODELSpecifies the source model. This can be a supported model repository identifier such as an OpenVINO model hosted on Hugging Face.
--model_repository_path PATHSpecifies where downloaded models should be stored locally.
--model_name NAMESpecifies the name OVMS exposes through its API.
--task text_generationTells OVMS that the model should be served using the text generation pipeline.
--target_device NPURuns inference on the Intel NPU. Other OpenVINO targets can include CPU, GPU, AUTO, and device-specific combinations.
--tool_parser hermes3Enables parsing model-generated Hermes-style tool calls. This is important when using the local model with agentic frontends that expect tool calling.
--max_prompt_len 16384Configures the maximum prompt/context length accepted by the serving pipeline.
--cache_dir .ovcacheStores OpenVINO's compiled model cache in the specified directory so future startups can avoid recompiling everything from scratch.
Basic Example
ovms.exe --pull --rest_port 8000 --source_model OpenVINO/Qwen2-0.5B-int4-ov --model_repository_path ".\models" --task text_generationNPU Qwen3 Example
This is the configuration I use for Qwen3 8B on the Intel NPU:
ovms.exe ^
--rest_port 8000 ^
--model_repository_path "%USERPROFILE%\Downloads\models" ^
--source_model OpenVINO/Qwen3-8B-int4-cw-ov ^
--model_name Qwen3-8B ^
--task text_generation ^
--target_device NPU ^
--tool_parser hermes3 ^
--max_prompt_len 16384 ^
--plugin_config "{\"NPUW_LLM_PREFILL_ATTENTION_HINT\":\"PYRAMID\"}" ^
--cache_dir .ovcacheTo verify that OVMS loaded the model, run:
curl http://localhost:8000/v3/modelsYou should see a model named:
Qwen3-8BOVMS also exposes OpenAI-compatible endpoints under:
http://localhost:8000/v1NoLlama
NoLlama provides another way to run OpenVINO-backed LLMs while exposing an OpenAI-compatible API.
Prerequisites
- Python 3.X
- OpenVINO libraries
NoLlama's installation script should install the required Python packages using pip.
- Intel NPU drivers if you want to run the model on the NPU.
Setup
Clone the repository:
git clone https://github.com/aweussom/NoLlama.gitRepository:
https://github.com/aweussom/NoLlamaReleases:
https://github.com/aweussom/NoLlama/releasesRun the installation script:
.\install.ps1Start NoLlama:
.\start.ps1
Verify that the server is running:
curl http://localhost:8000/v1/modelsNoLlama exposes an OpenAI-compatible server at:
http://localhost:8000/v1Frontend AI Assistants
Once an OpenAI-compatible backend is running, different frontend applications can connect to it.
OpenWork
OpenWork uses OpenCode internally and can directly use an OpenAI-compatible provider.
OpenWork With OVMS
In the OVMS folder, run
setupvars.bat.Start Qwen3 8B:
ovms.exe ^ --rest_port 8000 ^ --model_repository_path "%USERPROFILE%\Downloads\models" ^ --source_model OpenVINO/Qwen3-8B-int4-cw-ov ^ --model_name Qwen3-8B ^ --task text_generation ^ --target_device NPU ^ --tool_parser hermes3 ^ --max_prompt_len 16384 ^ --plugin_config "{\"NPUW_LLM_PREFILL_ATTENTION_HINT\":\"PYRAMID\"}" ^ --cache_dir .ovcacheCheck that OVMS is running and that the model is named Qwen3-8B:
curl http://localhost:8000/v3/modelsOpen your OpenWork workspace directory.
The default is typically:
%USERPROFILE%\OpenWork ChatReplace the contents of
opencode.jsoncwith:{ "$schema": "https://opencode.ai/config.json", "disabled_providers": [ "opencode" ], "provider": { "ovms": { "npm": "@ai-sdk/openai-compatible", "name": "OVMS Local", "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "ovms-local" }, "models": { "Qwen3-8B": { "name": "Qwen3 8B int4 (OVMS NPU)", "tool_call": true, "reasoning": true, "limit": { "context": 16384, "output": 4096 } } } } }, "model": "ovms/Qwen3-8B" }Restart OpenWork.
Open the model picker and select:
Qwen3 8B int4 (OVMS NPU)The first response can be considerably slower because OpenVINO may still be compiling and caching portions of the model.
OpenWork With NoLlama
Start NoLlama:
.\start.ps1Find the model ID exposed by NoLlama:
curl http://localhost:8000/v1/modelsConfigure OpenWork using the same OpenAI-compatible provider, replacing the model ID with the ID returned by NoLlama:
{ "$schema": "https://opencode.ai/config.json", "disabled_providers": [ "opencode" ], "provider": { "nollama": { "npm": "@ai-sdk/openai-compatible", "name": "NoLlama Local", "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "nollama-local" }, "models": { "YOUR-MODEL-ID": { "name": "NoLlama Local Model", "tool_call": true, "reasoning": true, "limit": { "context": 16384, "output": 4096 } } } } }, "model": "nollama/YOUR-MODEL-ID" }
OpenWork Troubleshooting
- HTTP 404
Make sure the model name in
opencode.jsoncexactly matches the model exposed by the backend.For OVMS, the OpenWork model ID must match
--model_name, including capitalization. - Model does not appear
Completely restart OpenWork after changing
opencode.jsonc. - Tools are not working correctly
Make sure OVMS was started with a tool parser compatible with the model, such as:
--tool_parser hermes3
Claude Code
Claude Code is slightly different from OpenWork.
OVMS and NoLlama expose an OpenAI-compatible API. Claude Code, however, normally communicates using Anthropic's Messages API.
Because the protocols are different, Claude Code cannot simply be pointed directly at:
http://localhost:8000/v1Instead, I use Claude Code Router as a local compatibility layer:
Claude Code
|
| Anthropic Messages API
v
Claude Code Router
http://127.0.0.1:3456
|
| OpenAI-compatible Chat API
v
OVMS / NoLlama
http://127.0.0.1:8000/v1
|
v
Qwen
|
v
Intel NPUPrerequisites
- Node.js 22 or newer
- Claude Code
- Claude Code Router
- A running OVMS or NoLlama server
Install Claude Code
Check your Node.js version:
node --versionInstall Claude Code:
npm install -g @anthropic-ai/claude-codeInstall Claude Code Router:
npm install -g @musistudio/claude-code-routerClaude Code With OVMS
Start OVMS using the same Qwen3 configuration:
ovms.exe ^ --rest_port 8000 ^ --model_repository_path "%USERPROFILE%\Downloads\models" ^ --source_model OpenVINO/Qwen3-8B-int4-cw-ov ^ --model_name Qwen3-8B ^ --task text_generation ^ --target_device NPU ^ --tool_parser hermes3 ^ --max_prompt_len 16384 ^ --plugin_config "{\"NPUW_LLM_PREFILL_ATTENTION_HINT\":\"PYRAMID\"}" ^ --cache_dir .ovcacheMake sure the model loaded:
curl http://localhost:8000/v3/modelsStart the Claude Code Router configuration interface:
ccr uiOpen the Claude Code Router configuration UI if it did not open automatically:
http://127.0.0.1:3458Go to the provider configuration and add a custom provider for OVMS.
Name: OVMS Local API Endpoint: http://127.0.0.1:8000/v1 API Key: ovms-local Protocol: OpenAI Chat Model: Qwen3-8BIf Claude Code Router automatically detects the wrong protocol, disable automatic protocol detection and explicitly select:
OpenAI ChatRun the provider connection test.
A successful response means Claude Code Router can communicate with OVMS.
Save the provider.
In Claude Code Router, create a local API/client key if your configuration requires one.
This does not need to be a real Anthropic API key. It is simply used by the local router.
Start the Claude Code Router server.
Its local gateway normally runs at:
http://127.0.0.1:3456Create or select a Claude Code profile that uses the OVMS Local provider and
Qwen3-8B.Launch Claude Code through Claude Code Router.
ccr codeSend a test message.
The request path should now be:
Claude Code -> Claude Code Router -> OVMS -> Qwen3-8B -> Intel NPU
Claude Code With NoLlama
NoLlama uses the same basic Claude Code Router configuration because it also exposes an OpenAI-compatible server.
Start NoLlama:
.\start.ps1Check which model IDs NoLlama exposes:
curl http://localhost:8000/v1/modelsAdd another provider in Claude Code Router:
Name: NoLlama Local API Endpoint: http://127.0.0.1:8000/v1 API Key: nollama-local Protocol: OpenAI Chat Model: YOUR-MODEL-IDReplace
YOUR-MODEL-IDwith the exact model ID returned by:curl http://localhost:8000/v1/modelsTest the provider and save it.
Select the NoLlama provider/model in your Claude Code Router profile.
Launch Claude Code through the router:
ccr code
Important: Claude Code vs Claude Desktop
This setup is specifically for Claude Code.
The normal Claude website or Claude Desktop chat application does not provide a setting that simply replaces Anthropic's Claude model with an arbitrary OpenAI-compatible local LLM.
Claude Desktop can use local MCP servers for tools and external information, but that is different from replacing Claude itself with Qwen running through OVMS.
Claude Code Troubleshooting
- Claude Code still contacts Anthropic
Make sure Claude Code was launched through Claude Code Router rather than by running Claude normally.
ccr code - HTTP 404 from OVMS
The router's model name must exactly match the OVMS model name.
For this example:
--model_name Qwen3-8B Model in Claude Code Router: Qwen3-8B - Normal chat works but tools fail
This normally means the base model request is working but tool call serialization or parsing is not.
Make sure OVMS was started with the appropriate tool parser:
--tool_parser hermes3 - Very slow first response
The first request after starting OVMS may require OpenVINO to compile the model for the NPU. The compiled model cache helps reduce this on later launches.
--cache_dir .ovcache
Final Setup
The entire local setup ends up looking roughly like this:
OpenWork
OpenWork
|
| OpenAI-compatible API
v
OVMS / NoLlama
|
v
OpenVINO
|
v
Intel AI Boost NPUClaude Code
Claude Code
|
| Anthropic Messages API
v
Claude Code Router
|
| OpenAI-compatible API
v
OVMS / NoLlama
|
v
OpenVINO
|
v
Intel AI Boost NPUThis gives me a small local model for background tasks, basic coding work, tool calls, and agent workflows without burning Claude tokens or occupying the laptop's RTX GPU.