Home | Markdown | Gemini | Microblog | Street photography | Wishlist

Running LLMs by myself - Part 2: The pi coding agent



Published at 2026-10-07T22:16:10+03:00

This is the second blog post about running LLMs by myself. The first part was about the model side: renting Hyperstack VMs, vLLM, how the inference works and what it all costs. This part is about the other half, the pi coding agent I actually type into.

2026-10-06 Running LLMs by myself - Part 1: Hyperstack and vLLM
2026-10-07 Running LLMs by myself - Part 2: The pi coding agent (You are currently reading this)

The hypr logo

Table of Contents




LLM vs harness



I say "the model edited my file" all the time, but that's not really what happens. Two different programs are involved:


Claude Code, Codex, OpenCode and Cursor are harnesses, too. They usually come bundled with their vendor's model, which blurs the line. With a self-hosted setup the line is very visible: the model is on a VM in Canada, the harness is on my laptop, and the only thing between them is HTTP over WireGuard.

So the model decides what to do, and the harness does it. That's why I can swap the model under pi with one keystroke without changing anything else. It's also why the same model can feel smart in one harness and dumb in another. The harness decides what the model gets to see, and that has a huge effect on how good "the LLM" seems. More on that once the system prompt, tools and skills are explained.

How pi talks to vLLM



Pi speaks the OpenAI chat completions API, and vLLM serves exactly that on :11434, so pi points straight at the VM over the tunnel, with no translation proxy in between.

The repo ships a pi/ directory that I symlink to ~/.pi. It defines providers in models.json, one per VM, plus a single-VM variant:


Every preset from the TOML configs is registered under its provider, so after a model switch I can just tell pi to use the new model ID, or hit Ctrl+L in the TUI to switch models mid-session without restarting.

Fish abbreviations keep the day-to-day short:

abbr pi-hyperstack-coder  pi --model hyperstack1/Qwen/Qwen3.8-27B-FP8
abbr pi-hyperstack-gemma4 pi --model hyperstack2/cyankiwi/gemma-4-31B-it-AWQ-4bit

My standard setup is a tmux session with one pi per pane: pi-hyperstack-coder on Qwen3.8 in pane 0, pi-hyperstack-gemma4 in pane 1, each working on a different project against its own VM. When one model gets stuck on a task, I hand the same problem to the other pane and compare.

Here is what a session looks like in practice. I asked Qwen3.8 on VM1 to add a version command to the REPL of gt, and then to include the Go runtime version in its output. The screenshot shows the Go diff for the tests, followed by the test run. The footer shows the model and how much of the 262K context the session has used so far:

pi with Qwen3.8 27B editing Go code in the gt project, followed by the test run

And one from an earlier session: diff on top, the model's reasoning in the middle, shell output at the bottom:

pi coding agent mid-task: diff, reasoning, and shell in one TUI

The system prompt and harness overhead



Pi puts a system prompt in front of the conversation. It tells the model what it's there for, how to use the tools, and which project rules apply.

This is the start of the built-in system prompt from the pi version I used for these sessions, trimmed a little. Current upstream pi builds it in named sections, so the exact layout below is a snapshot of my setup:

You are an expert coding assistant operating inside pi, a coding agent
harness. You help users by reading files, executing commands, editing
code, and writing new files.

Available tools:
- read: Read file contents
- bash: Execute bash commands (ls, grep, find, etc.)
- edit: Make precise file edits with exact text replacement, including
  multiple disjoint edits in one call
- write: Create or overwrite files

Guidelines:
- Use read to examine files instead of cat or sed.
- Use edit for precise changes (edits[].oldText must match exactly)
- Use write only for new files or complete rewrites.
- Be concise in your responses
- Show file paths clearly when working with files

In that version, the project context (AGENTS.md or CLAUDE.md) came next, followed by the skills, then the date and working directory:

<project_context>
Project-specific instructions and guidelines:
<project_instructions path="/home/paul/git/hypr/AGENTS.md">
...
</project_instructions>
</project_context>

The following skills provide specialized instructions for specific tasks.
Use the read tool to load a skill's file when the task matches its description.

<available_skills>
  <skill>
    <name>solid-principles</name>
    <description>This skill should be used when the user asks to "check SOLID violations", "audit class design", ...</description>
    <location>/home/paul/.agents/skills/solid-principles/SKILL.md</location>
  </skill>
  ...
</available_skills>

Current date: 2026-09-30
Current working directory: /home/paul/git/hypr

Putting the stable instructions first gives prefix caching more to reuse. The layout has changed since these sessions; current upstream's system-prompt builder no longer adds the date shown here.

Pi's system-prompt builder

You never type it, but it's resent with every request and takes up KV cache like everything else. So do all the tool definitions (bash, read, edit, web_search, ...), skill descriptions (more on skills below), and project instructions such as an AGENTS.md. With dozens of tools, that's thousands of tokens before you've typed a word. While those definitions stay the same and the cache blocks are still there, vLLM can reuse that prefix as explained in part 1.

How tool calling works



The model doesn't run anything itself. The harness sends it a list of tool schemas (name, description, JSON arguments) along with the prompt. When the model wants to act, it outputs a structured call such as read_file {"path": "main.go"} instead of prose. The harness runs the tool, appends the result to the conversation, and asks the model again. That repeats until the task is done.

A malformed call or the wrong tool choice can stop that loop.

One agent turn: the tool-calling loop between pi and vLLM

Tool calling on the wire



So what does this look like on the wire? Since vLLM speaks the OpenAI chat completions API, the whole agent loop is ordinary HTTP. Here's one turn, trimmed down. First, the harness (pi) sends the conversation plus the tool schemas:

POST http://hyperstack1.wg1:11434/v1/chat/completions
{
  "model": "Qwen/Qwen3.8-27B-FP8",
  "messages": [
    {"role": "system", "content": "You are a coding agent. Use the tools..."},
    {"role": "user",   "content": "What does main.go do?"}
  ],
  "tools": [{
    "type": "function",
    "function": {
      "name": "read_file",
      "description": "Read a file from the project",
      "parameters": {
        "type": "object",
        "properties": {"path": {"type": "string"}},
        "required": ["path"]
      }
    }
  }]
}

The model doesn't answer the question yet. It answers with a tool call instead of text. This is the real (trimmed) response from Qwen3.8 on VM1:

{
  "role": "assistant",
  "content": null,
  "tool_calls": [{
    "id": "chatcmpl-tool-996eae953d7c06fa",
    "type": "function",
    "function": {"name": "read_file", "arguments": "{\"path\": \"main.go\"}"}
  }]
}

The API returns arguments as a JSON string. With the qwen3_coder parser that hypr sets for Qwen, the model generates tagged function and parameter text, and vLLM turns it into that JSON. A malformed call can still break parsing. The harness validates the arguments, runs the tool locally, and sends everything back with the result appended:

"messages": [
  {"role": "system",    "content": "You are a coding agent. Use the tools..."},
  {"role": "user",      "content": "What does main.go do?"},
  {"role": "assistant", "tool_calls": [{"id": "chatcmpl-tool-996e...", ... "read_file" ...}]},
  {"role": "tool",      "tool_call_id": "chatcmpl-tool-996e...", "content": "package main\n\nfunc main() {..."}
]

By the way, vLLM reported 292 prompt tokens for that first request, with just one short system prompt and one tool. pi's real system prompt with all its tools is a lot bigger.

Now the model has the file contents in its context and can answer in plain text (or request another tool call, and the loop goes on). Two things I found interesting here. In this text-only workflow, the model sees text going in and text coming out, so "calling a tool" is just a special output format it was trained to produce. And every round trip resends the whole history, including all tool results, so the context (and the KV cache) grows with every step. That is why agentic work is so prefix-cache-heavy.

For automatic tool selection in my setup, vLLM needs --enable-auto-tool-choice and a matching --tool-call-parser, here qwen3_coder. hypr sets these per preset. A wrong parser can leave raw text where the agent expects a tool call. Named and required tool calling also have structured-output paths that work without enabling automatic tool selection.

vLLM tool calling

Skills, commands and MCP servers



Pi has prompt templates and skills. I initially thought of them as manual versus automatic prompts, but there is some overlap:


Pi lists the skills available for automatic selection in the system prompt. A skill with disable-model-invocation: true stays out of that list; I have to invoke it myself. For the advertised skills, loading works in three steps:


Progressive skill loading for automatically available skills
Pi skills and explicit invocation

Those advertised descriptions cost context on every request, even for skills I never use. Levels 2 and 3 cost nothing until they're loaded. But once the model has read a SKILL.md or a reference file, it's a tool result in the conversation, and it stays in the context (and the KV cache) for the rest of the session, until a compaction or a /handoff throws it out. My solid-principles skill is a good example: its SKILL.md is ~1.2K tokens, and the one reference file a single-principle check needs (say srp.md) is another ~1K. The other four reference files (~5K tokens) stay on disk unless a task asks for them. Loading everything up front would cost ~7K tokens, which is almost a quarter of a 32K preset.

I have 45 skills and 21 commands in my pi setup. The skill descriptions alone are ~17 KB of text. With names, paths and tags, the whole list is about 25 KB, roughly 6K tokens in every request. On the 262K daily driver, that's fine. On a 32K preset, it's almost a fifth of the context gone before I've typed a word.

The bigger downside of too many skills isn't even the tokens, though. The model has to pick the right skill from the list, and with many similar descriptions, it picks the wrong one or none at all. Big frontier models handle that pretty well. Smaller self-hosted models get confused much more easily. So fewer, clearly distinct skills work better, especially with local models.

MCP (Model Context Protocol) connects a harness to tools and resources from another process or service, say a database or a browser. Current upstream pi supports MCP over stdio and HTTP. It can expose tools directly, load their definitions through tool search, or let the model call them through code. Connecting a server with dozens of tools doesn't have to put all their schemas into every request.

I don't use MCP in this setup. CLI tools and skills cover what I need so far. Mario Zechner's earlier post explains that approach, though its description of pi predates the current MCP support.

Pi's current MCP support
What if you don't need MCP? (Mario Zechner)

Why the harness makes the model look smart (or dumb)



Pi decides which project details the model gets to see and which tools it can use. A few things matter in practice:


The screenshot above is a good example. Qwen3.8 didn't just write the version command. It looked at how the other builtins are registered, added a matching help entry, and ran go vet and the tests before saying it was done. The model did the thinking, but it could only do that because pi gave it file reading, editing and a shell, and fed the test output back into the context.

This also matters for benchmarks. Scores like SWE-bench are measured with a specific agent setup around the model. Put the same weights into a different harness, and you get different results, sometimes a lot better or worse. So when a self-hosted model disappoints, it's worth checking the harness side (tools, context, parser, settings) before blaming the weights.

The extensions



Pi has no built-in plan mode or sub-agents, and it doesn't ask for approval before every tool call. Current upstream does have project-trust prompts before loading project extensions and other resources.

Pi's project trust and tool permissions

I use TypeScript extensions from the hypr repo for the rest. These are the ones I use daily:


The long tail is smaller: nemotron-tool-repair fixes the malformed tool calls the Nemotron models occasionally emit, and prompt-history quietly records the last 500 prompts.

For the record: handoff, inline-bash, session-name, and reload-runtime are upstream pi examples installed locally; the rest are my own (created with the help of LLMs).

What the extensions cost



Tools exposed directly to the model add their definitions to its input, just as advertised skills add their descriptions. Deferred tools only add their full definitions when loaded. I haven't measured the exact token bill of the full extension set, but it isn't free. On the 27B FP8 daily driver the full set is fine. On the smaller AWQ presets with 32K context, I'd start with fewer tools and add them when needed, since fewer tools also means fewer chances for a small model to pick the wrong one or mangle the arguments. That's a hypothesis, though. I haven't measured it.

Tool calling in practice



I haven't compared tool-call reliability systematically across the preset list, so no failure-rate table from me. What I know from daily use: the Nemotron models occasionally emit malformed tool calls, which is why nemotron-tool-repair exists. It patches the broken calls so the session continues instead of stalling.

What I trust it with (and what I don't)



The gt calculator was the proof case, a real project built almost entirely on this stack. Day to day I also use it for ops babysitting via /loop and /watch: check the VM, watch a build, poke me when something smells wrong. That is work I trust it with.

Vendor SWE-bench numbers are not my session success rate. Hit-and-miss still happens. When a turn starts looping or the model gets lost in its own plan, I bounce the hard bit to a hosted frontier model and bring the answer back.

It's good enough that I built a real project on it and keep using it every day. But I still have to babysit it.

Wrapping up



That was the harness side of my setup. The pi extensions, together with the provisioner and the model presets from part 1, are all here:

hypr on GitHub

Read the previous post of this series:

Running LLMs by myself - Part 1: Hyperstack and vLLM

Other related posts:

2026-10-07 Running LLMs by myself - Part 2: The pi coding agent (You are currently reading this)
2026-10-06 Running LLMs by myself - Part 1: Hyperstack and vLLM
2026-06-01 gt calculator - a calculator built with local LLMs
2025-08-05 Local LLM for Coding with Ollama on macOS

E-Mail your comments to paul@nospam.buetow.org :-)

Back to the main site