Files
SnarfCode/reference_guide.md

5.3 KiB

Reference Guide: Open WebUI & Hailo-10H NPU Stack

This document serves as a complete technical reference for the architecture, service definitions, API schemas, and deployment designs implemented to connect Open WebUI (on Pi 4) with the Hailo-10H NPU HAT+ (on Pi 5).


1. System Architecture Layout

graph TD
    subgraph Host_Pi4 ["Debian 13 Pi 4 (Frontend Host)"]
        UI["Open WebUI (Native Python)<br>Port 8080"]
        Bridge["Pipelines Service (uvicorn)<br>Port 9099<br>Module: hailo_pipeline.py"]
    end

    subgraph Host_Pi5 ["Raspberry Pi 5 (NPU Host)"]
        NPU_Server["hailo-ollama Server (oatpp)<br>Port 8000"]
        NPU_HW["Hailo-10H NPU HAT+<br>(8GB VRAM / 40 TOPS)"]
    end

    %% Communications Flow
    User -->|Browser HTTP| UI
    UI -->|OpenAI Chat Request| Bridge
    Bridge -->|Ollama Chat JSON Payload| NPU_Server
    NPU_Server -->|Model Loading & Graph Execution| NPU_HW

2. Host 1: Raspberry Pi 5 (NPU Inference Engine)

The Pi 5 hosts the physical PCIe Hailo HAT+ hardware and runs the bare-metal, native C++ compiled server hailo-ollama powered by the Oatpp framework.

Service Definition: /etc/systemd/system/hailo-ollama.service

[Unit]
Description=Hailo-Ollama NPU Server
After=network-online.target

[Service]
User=root
Group=root
ExecStart=/usr/bin/hailo-ollama serve
Restart=always
RestartSec=3
Environment="OLLAMA_HOST=0.0.0.0:8000"
Environment="HAILO_MONITOR=1"

[Install]
WantedBy=multi-user.target

Critical Paths & Directories

  • Binary Path: /usr/bin/hailo-ollama
  • Configuration: /etc/xdg/hailo-ollama/hailo-ollama.json
  • Model Manifests (JSON Templates): /usr/share/hailo-ollama/models/manifests/
    • Contains system templates, repetition penalties, stop tokens, and target HEF hashes for models.
  • Model Blob Storage (Downloaded HEF Files): /usr/share/hailo-ollama/models/blob/
    • Downloaded model binaries compiled for the Hailo NPU. Named by their SHA-256 hash.

3. Host 2: Raspberry Pi 4 (WebUI & Pipelines Bridge)

The Pi 4 runs a native Python 3.11 environment hosting Open WebUI and the Pipelines server. The custom Pipeline script bridges the API mismatch and translates OpenAI requests to Oatpp.

Service Definition: /etc/systemd/system/pipelines.service

[Unit]
Description=Open WebUI Pipelines Bridge
After=network.target

[Service]
User=root
Group=root
WorkingDirectory=/root/open-webui/pipelines
Environment="PATH=/root/open-webui/venv/bin:/usr/local/bin:/usr/bin:/bin"
Environment="PIPELINES_API_KEY=hailo-bridge"
ExecStart=/root/open-webui/venv/bin/uvicorn main:app --host 0.0.0.0 --port 9099 --workers 1
Restart=always
RestartSec=3

[Install]
WantedBy=multi-user.target

Script Location: /root/open-webui/pipelines/pipelines/hailo_pipeline.py

This script implements:

  1. Manifold Pipeline Spec: Exposes sub-models cleanly to Open WebUI's dropdown.
  2. Namespace Stripping: Cleans model identifiers prepended by WebUI.
  3. Response Key Hierarchy: Safely extracts inference text from different key structures (Ollama nested, direct content/response keys, OpenAI choices).
  4. Graceful Streaming Loops: Handles SSE line-by-line buffers and yields tokens before evaluating completion flags.

4. API & Communication Schemas (Oatpp)

The hailo-ollama service mimics the Ollama REST API. However, it requires explicit model preparation.

4.1. Pull Model Endpoint (Downloads weights to local NPU blob store)

  • URL: POST http://140.44.4.7:8000/api/pull
  • Request Payload:
    {
      "model": "llama3.2:3b"
    }
    
  • Behavior: Downloads the model over the internet and compiles/saves the hardware HEF binary to /usr/share/hailo-ollama/models/blob/. If the model is not pulled, /api/chat returns a 200 OK with "error": "model not found".

4.2. Chat Inference Endpoint

  • URL: POST http://140.44.4.7:8000/api/chat
  • Request Payload:
    {
      "model": "llama3.2:3b",
      "messages": [
        {"role": "user", "content": "Hi"}
      ],
      "stream": true
    }
    

Non-Streaming Response ("stream": false)

{
  "model": "llama3.2:3b",
  "created_at": "2026-07-09T14:00:30.299821983Z",
  "message": {
    "role": "assistant",
    "content": "Hello! How can I help you today?"
  },
  "done": true,
  "done_reason": "stop",
  "total_duration": 22737674048,
  "eval_count": 58
}

Streaming Response Chunks ("stream": true)

Sent as a series of JSON strings separated by newlines:

{"model":"llama3.2:3b","message":{"role":"assistant","content":"Why"},"done":false}
{"model":"llama3.2:3b","message":{"role":"assistant","content":" couldn"},"done":false}
...
{"model":"llama3.2:3b","message":{"role":"assistant","content":""},"done":true,"done_reason":"stop"}

5. Operations & Troubleshooting Commands

Tail Service Logs

# On Pi 4 (WebUI / Pipelines Bridge)
journalctl -u open-webui -f
journalctl -u pipelines -f

# On Pi 5 (NPU Engine)
journalctl -u hailo-ollama -f

Restart Services

# On Pi 4
systemctl restart pipelines
systemctl restart open-webui

# On Pi 5
systemctl restart hailo-ollama

Clean Stuck NPU Resource Locks

If the C++ server hangs while compiling or loading a model, it can lock /dev/hailo0. Use kill -9 to force-release the NPU:

# On Pi 5
sudo kill -9 $(pgrep hailo-ollama)
sudo systemctl restart hailo-ollama