An MCP server that lets an agent write WebGPU shaders

mcp, webgpu, tsl, three-js, agents, architecture

You cannot render WebGPU in a Node process. So an MCP server that wants to let an agent write shaders has a problem before it starts: the thing that compiles the shader and the thing the agent talks to cannot live in the same place.

Shader Lab's MCP server solves it by owning nothing. It is a relay.

MCP client (Claude Code) ⇄ stdio ⇄ shader-lab-mcp (Bun process)
                                       ⇅ WebSocket (127.0.0.1:7420)
                                 editor tab (?agent=1)

The server speaks MCP over stdio and hosts a loopback-only WebSocket bridge. The editor connects to that bridge when you open it with ?agent=1. Every tool call is relayed into the tab and executed through the editor's normal store actions — which means everything the agent does lands in the undo history, and Cmd+Z works on it.

The renderer stays where the GPU is. The server stays where the agent is. Nothing has to move.

The loop is the product

The tools that create and reorder layers are the boring part. The one that matters is write_custom_shader, and what makes it work is not that it writes a file — it is what comes back:

Returns { compiled, error } — on failure, `error` is the exact
compiler/runtime message, so fix and retry.

That is the whole thing. An agent writing shader code without the compiler's exact output is guessing, and guessing at TSL produces confident nonsense. Hand back the verbatim error synchronously and the agent converges in two or three attempts, because it is doing what a person does: read the error, fix that line, recompile.

Then screenshot closes it. Compiling is not the same as being right, and a shader can compile perfectly into a black rectangle. The agent writes, compiles, reads the error or reads the pixels, and goes again.

Write → compile → exact diagnostic → fix → look at the pixels. Four steps, and the middle two are the ones people skip when they build this kind of tool.

No imports, on purpose

The contract for a sketch is deliberately small:

export const sketch = Fn(() => {
  // returns a TSL node
})

No import statements are allowed. Everything — all of three/tsl, the house utilities, time, and inputTexture when the shader is running in effect mode — is injected as a global.

That looks like a limitation and is actually the point. Module resolution is one of the most reliable ways to make a language model produce broken code: it will import a real function from a plausible-but-wrong path, and the failure arrives as a resolution error that has nothing to do with the shader it was trying to write. Delete the import statement from the language and that entire failure class stops existing. There is no dependency graph to get wrong because there are no dependencies.

effectMode picks the other axis: false generates imagery from scratch, true transforms whatever the layer stack below produces, sampled through inputTexture.

What it is not

Being precise, because "live compilation" covers a lot of ground:

  • Not incremental. The sketch recompiles as a unit. There is no dependency tracking and nothing is cached between attempts.
  • Not module hot-reload. There are no modules. See above.
  • Not batched. One tool call, one compile, one answer.

What it is: a synchronous compile against a live WebGPU renderer, with the real diagnostic returned to the caller inside a 15-second budget, in a tab that is already showing the result. That covers the case that matters — an agent iterating on a shader — without pretending to be a compiler toolchain.

The unglamorous parts that make it usable

Timeouts are per-operation, not global. 5s default, 15s for a compile, 30s for a screenshot. A shader compile is not a state read and should not share its budget.

The bridge is loopback-only, and the origin list is explicit — localhost, plus eng.basement.studio and *.vercel.app previews, extensible by env var. A deployed editor tab connects back to your machine, so the deployment is never in the path. The server runs locally and the bridge never leaves it.

One tab at a time. The bridge holds a single connection. Two editors would mean two renderers disagreeing about which is authoritative, and there is no good answer to that, so it is not allowed.

The API reference is generated. get_shader_api_reference is built from a generation script rather than hand-written, so the surface described to the agent cannot drift from the surface that exists. Agent-facing documentation that goes stale is worse than none — it teaches the model to call things that were removed.

Errors are addressed to whoever can fix them. Not "connection refused" but open the editor with ?agent=1 appended and try again, and for a port clash, the port number and the env var that changes it. The reader of an MCP error is usually an agent that will act on it immediately, so the message should say what to do rather than what went wrong.

The takeaway

If you are giving an agent control of something visual, the temptation is to spend the effort on the API surface — more tools, more parameters, more coverage. That is the wrong end. The surface barely matters. What determines whether the agent can actually work is whether failure comes back specifically and fast, and whether it can see the result afterwards.

Exact compiler output and a screenshot. Everything else is layers and plumbing.