Title

LLMs on an Intel NPU

How I got Qwen3 8B running on my laptop's Intel AI Boost NPU with OpenVINO Model Server (or NoLlama), then hooked it into OpenWork and Claude Code, so small tasks stop burning Claude tokens and the GPU stays free.

Setup

Abstract

I have an OMEN MAX Gaming Laptop 16 and recently I have been working on a lot of side projects in my free time. I've also been using Claude a lot since I got this laptop to become Peter Thiel's most feared 10x engineer.

Unfortunately, Claude has the habit of spawning more workflows, and that just absolutely DRAINS my token usage and limits.

Eventually I wanted to use a locally hosted model for smaller tasks and found out my laptop had an Intel AI Boost NPU that I could use instead of my GPU. This lets me run smaller LLM workloads without tying up the GPU while I am doing other work.

Then I realized there was very little documentation explaining how to actually use this POS hardware chip for an LLM. After a lot of experimenting, I managed to get Qwen running on the NPU through OpenVINO.

This page documents the setup I ended up using for OpenVINO Model Server, NoLlama, OpenWork, and Claude Code.


Backend Model Server

The frontend AI application and the model server are separate pieces.

Frontend
   |
   | OpenAI / Anthropic compatible API
   v
Model Server
   |
   v
OpenVINO
   |
   v
Intel NPU

For the backend, I have mainly used OpenVINO Model Server and NoLlama.

OpenVINO Model Server

Prerequisites

  • Python 3.12.X - Optional

    OpenVINO Model Server has Windows packages with Python enabled and Python disabled. Some additional features and model workflows may require the Python-enabled package.

  • Intel NPU Drivers

    Download and install the latest Intel NPU drivers from Intel:

    Intel NPU Driver for Windows

Installation

  1. Download OpenVINO Model Server from the official repository:

    https://github.com/openvinotoolkit/model_server
  2. Extract OVMS somewhere on your computer.

    For example:

    C:\Users\YOUR_USERNAME\Downloads\ovms
  3. Open Command Prompt inside the OVMS directory and initialize the OpenVINO environment:

    setupvars.bat

Useful OVMS Command Line Arguments

  • --rest_port 8000

    Starts the HTTP REST API on port 8000. The OpenAI-compatible API is exposed through this server.

  • --rest_bind_address 127.0.0.1

    Binds OVMS only to the local machine instead of exposing it to other devices on the network.

  • --rest_workers N

    Controls the number of HTTP REST worker threads.

  • --port 9000

    Enables the gRPC server on port 9000.

  • --grpc_workers N

    Controls the number of gRPC worker threads.

  • --pull

    Downloads the requested model before starting the server.

  • --source_model MODEL

    Specifies the source model. This can be a supported model repository identifier such as an OpenVINO model hosted on Hugging Face.

  • --model_repository_path PATH

    Specifies where downloaded models should be stored locally.

  • --model_name NAME

    Specifies the name OVMS exposes through its API.

  • --task text_generation

    Tells OVMS that the model should be served using the text generation pipeline.

  • --target_device NPU

    Runs inference on the Intel NPU. Other OpenVINO targets can include CPU, GPU, AUTO, and device-specific combinations.

  • --tool_parser hermes3

    Enables parsing model-generated Hermes-style tool calls. This is important when using the local model with agentic frontends that expect tool calling.

  • --max_prompt_len 16384

    Configures the maximum prompt/context length accepted by the serving pipeline.

  • --cache_dir .ovcache

    Stores OpenVINO's compiled model cache in the specified directory so future startups can avoid recompiling everything from scratch.

Basic Example

ovms.exe --pull --rest_port 8000 --source_model OpenVINO/Qwen2-0.5B-int4-ov --model_repository_path ".\models" --task text_generation

NPU Qwen3 Example

This is the configuration I use for Qwen3 8B on the Intel NPU:

ovms.exe ^
  --rest_port 8000 ^
  --model_repository_path "%USERPROFILE%\Downloads\models" ^
  --source_model OpenVINO/Qwen3-8B-int4-cw-ov ^
  --model_name Qwen3-8B ^
  --task text_generation ^
  --target_device NPU ^
  --tool_parser hermes3 ^
  --max_prompt_len 16384 ^
  --plugin_config "{\"NPUW_LLM_PREFILL_ATTENTION_HINT\":\"PYRAMID\"}" ^
  --cache_dir .ovcache

To verify that OVMS loaded the model, run:

curl http://localhost:8000/v3/models

You should see a model named:

Qwen3-8B

OVMS also exposes OpenAI-compatible endpoints under:

http://localhost:8000/v1

NoLlama

NoLlama provides another way to run OpenVINO-backed LLMs while exposing an OpenAI-compatible API.

Prerequisites

  • Python 3.X
  • OpenVINO libraries

    NoLlama's installation script should install the required Python packages using pip.

  • Intel NPU drivers if you want to run the model on the NPU.

Setup

  1. Clone the repository:

    git clone https://github.com/aweussom/NoLlama.git

    Repository:

    https://github.com/aweussom/NoLlama

    Releases:

    https://github.com/aweussom/NoLlama/releases
  2. Run the installation script:

    .\install.ps1
  3. Start NoLlama:

    .\start.ps1

Verify that the server is running:

curl http://localhost:8000/v1/models

NoLlama exposes an OpenAI-compatible server at:

http://localhost:8000/v1

Frontend AI Assistants

Once an OpenAI-compatible backend is running, different frontend applications can connect to it.

OpenWork

OpenWork uses OpenCode internally and can directly use an OpenAI-compatible provider.

OpenWork With OVMS

  1. In the OVMS folder, run setupvars.bat.

  2. Start Qwen3 8B:

    ovms.exe ^
      --rest_port 8000 ^
      --model_repository_path "%USERPROFILE%\Downloads\models" ^
      --source_model OpenVINO/Qwen3-8B-int4-cw-ov ^
      --model_name Qwen3-8B ^
      --task text_generation ^
      --target_device NPU ^
      --tool_parser hermes3 ^
      --max_prompt_len 16384 ^
      --plugin_config "{\"NPUW_LLM_PREFILL_ATTENTION_HINT\":\"PYRAMID\"}" ^
      --cache_dir .ovcache
  3. Check that OVMS is running and that the model is named Qwen3-8B:

    curl http://localhost:8000/v3/models
  4. Open your OpenWork workspace directory.

    The default is typically:

    %USERPROFILE%\OpenWork Chat
  5. Replace the contents of opencode.jsonc with:

    {
      "$schema": "https://opencode.ai/config.json",
      "disabled_providers": [
        "opencode"
      ],
      "provider": {
        "ovms": {
          "npm": "@ai-sdk/openai-compatible",
          "name": "OVMS Local",
          "options": {
            "baseURL": "http://localhost:8000/v1",
            "apiKey": "ovms-local"
          },
          "models": {
            "Qwen3-8B": {
              "name": "Qwen3 8B int4 (OVMS NPU)",
              "tool_call": true,
              "reasoning": true,
              "limit": {
                "context": 16384,
                "output": 4096
              }
            }
          }
        }
      },
      "model": "ovms/Qwen3-8B"
    }
  6. Restart OpenWork.

  7. Open the model picker and select:

    Qwen3 8B int4 (OVMS NPU)
  8. The first response can be considerably slower because OpenVINO may still be compiling and caching portions of the model.

OpenWork With NoLlama

  1. Start NoLlama:

    .\start.ps1
  2. Find the model ID exposed by NoLlama:

    curl http://localhost:8000/v1/models
  3. Configure OpenWork using the same OpenAI-compatible provider, replacing the model ID with the ID returned by NoLlama:

    {
      "$schema": "https://opencode.ai/config.json",
      "disabled_providers": [
        "opencode"
      ],
      "provider": {
        "nollama": {
          "npm": "@ai-sdk/openai-compatible",
          "name": "NoLlama Local",
          "options": {
            "baseURL": "http://localhost:8000/v1",
            "apiKey": "nollama-local"
          },
          "models": {
            "YOUR-MODEL-ID": {
              "name": "NoLlama Local Model",
              "tool_call": true,
              "reasoning": true,
              "limit": {
                "context": 16384,
                "output": 4096
              }
            }
          }
        }
      },
      "model": "nollama/YOUR-MODEL-ID"
    }

OpenWork Troubleshooting

  • HTTP 404

    Make sure the model name in opencode.jsonc exactly matches the model exposed by the backend.

    For OVMS, the OpenWork model ID must match --model_name, including capitalization.

  • Model does not appear

    Completely restart OpenWork after changing opencode.jsonc.

  • Tools are not working correctly

    Make sure OVMS was started with a tool parser compatible with the model, such as:

    --tool_parser hermes3

Claude Code

Claude Code is slightly different from OpenWork.

OVMS and NoLlama expose an OpenAI-compatible API. Claude Code, however, normally communicates using Anthropic's Messages API.

Because the protocols are different, Claude Code cannot simply be pointed directly at:

http://localhost:8000/v1

Instead, I use Claude Code Router as a local compatibility layer:

Claude Code
     |
     | Anthropic Messages API
     v
Claude Code Router
http://127.0.0.1:3456
     |
     | OpenAI-compatible Chat API
     v
OVMS / NoLlama
http://127.0.0.1:8000/v1
     |
     v
Qwen
     |
     v
Intel NPU

Prerequisites

  • Node.js 22 or newer
  • Claude Code
  • Claude Code Router
  • A running OVMS or NoLlama server

Install Claude Code

Check your Node.js version:

node --version

Install Claude Code:

npm install -g @anthropic-ai/claude-code

Install Claude Code Router:

npm install -g @musistudio/claude-code-router

Claude Code With OVMS

  1. Start OVMS using the same Qwen3 configuration:

    ovms.exe ^
      --rest_port 8000 ^
      --model_repository_path "%USERPROFILE%\Downloads\models" ^
      --source_model OpenVINO/Qwen3-8B-int4-cw-ov ^
      --model_name Qwen3-8B ^
      --task text_generation ^
      --target_device NPU ^
      --tool_parser hermes3 ^
      --max_prompt_len 16384 ^
      --plugin_config "{\"NPUW_LLM_PREFILL_ATTENTION_HINT\":\"PYRAMID\"}" ^
      --cache_dir .ovcache
  2. Make sure the model loaded:

    curl http://localhost:8000/v3/models
  3. Start the Claude Code Router configuration interface:

    ccr ui
  4. Open the Claude Code Router configuration UI if it did not open automatically:

    http://127.0.0.1:3458
  5. Go to the provider configuration and add a custom provider for OVMS.

    Name: OVMS Local
    
    API Endpoint:
    http://127.0.0.1:8000/v1
    
    API Key:
    ovms-local
    
    Protocol:
    OpenAI Chat
    
    Model:
    Qwen3-8B
  6. If Claude Code Router automatically detects the wrong protocol, disable automatic protocol detection and explicitly select:

    OpenAI Chat
  7. Run the provider connection test.

    A successful response means Claude Code Router can communicate with OVMS.

  8. Save the provider.

  9. In Claude Code Router, create a local API/client key if your configuration requires one.

    This does not need to be a real Anthropic API key. It is simply used by the local router.

  10. Start the Claude Code Router server.

    Its local gateway normally runs at:

    http://127.0.0.1:3456
  11. Create or select a Claude Code profile that uses the OVMS Local provider and Qwen3-8B.

  12. Launch Claude Code through Claude Code Router.

    ccr code
  13. Send a test message.

    The request path should now be:

    Claude Code
        ->
    Claude Code Router
        ->
    OVMS
        ->
    Qwen3-8B
        ->
    Intel NPU

Claude Code With NoLlama

NoLlama uses the same basic Claude Code Router configuration because it also exposes an OpenAI-compatible server.

  1. Start NoLlama:

    .\start.ps1
  2. Check which model IDs NoLlama exposes:

    curl http://localhost:8000/v1/models
  3. Add another provider in Claude Code Router:

    Name: NoLlama Local
    
    API Endpoint:
    http://127.0.0.1:8000/v1
    
    API Key:
    nollama-local
    
    Protocol:
    OpenAI Chat
    
    Model:
    YOUR-MODEL-ID
  4. Replace YOUR-MODEL-ID with the exact model ID returned by:

    curl http://localhost:8000/v1/models
  5. Test the provider and save it.

  6. Select the NoLlama provider/model in your Claude Code Router profile.

  7. Launch Claude Code through the router:

    ccr code

Important: Claude Code vs Claude Desktop

This setup is specifically for Claude Code.

The normal Claude website or Claude Desktop chat application does not provide a setting that simply replaces Anthropic's Claude model with an arbitrary OpenAI-compatible local LLM.

Claude Desktop can use local MCP servers for tools and external information, but that is different from replacing Claude itself with Qwen running through OVMS.

Claude Code Troubleshooting

  • Claude Code still contacts Anthropic

    Make sure Claude Code was launched through Claude Code Router rather than by running Claude normally.

    ccr code
  • HTTP 404 from OVMS

    The router's model name must exactly match the OVMS model name.

    For this example:

    --model_name Qwen3-8B
    
    Model in Claude Code Router:
    Qwen3-8B
  • Normal chat works but tools fail

    This normally means the base model request is working but tool call serialization or parsing is not.

    Make sure OVMS was started with the appropriate tool parser:

    --tool_parser hermes3
  • Very slow first response

    The first request after starting OVMS may require OpenVINO to compile the model for the NPU. The compiled model cache helps reduce this on later launches.

    --cache_dir .ovcache

Final Setup

The entire local setup ends up looking roughly like this:

OpenWork

OpenWork
   |
   | OpenAI-compatible API
   v
OVMS / NoLlama
   |
   v
OpenVINO
   |
   v
Intel AI Boost NPU

Claude Code

Claude Code
   |
   | Anthropic Messages API
   v
Claude Code Router
   |
   | OpenAI-compatible API
   v
OVMS / NoLlama
   |
   v
OpenVINO
   |
   v
Intel AI Boost NPU

This gives me a small local model for background tasks, basic coding work, tool calls, and agent workflows without burning Claude tokens or occupying the laptop's RTX GPU.