Lightweight Benchmark Server
for Agentic GPU Programming

Shared GPU capacity. Reliable evaluation.
Independent scaling.

Use cases

Kernel Evolution

Share a smaller GPU pool across agents. Exclusive access and optional workspace isolation help evaluations stay reliable.

Kernel Post-training

Scale evaluations across GPUs and nodes through a unified entry point. Grow agent workloads and compute capacity independently.

How KCoral works

Agents generate and revise code in their own workspaces. KCoral runs the evaluation and returns feedback.

Your workspace Agent Generate & revise
Evaluation service Schedule & execute
Shared hardware GPUs Check · Benchmark · Profile
Selected results, execution status, and logs flow back to the agent.

KCoral Protocol

Each request is a self-contained program over HTTP. Compose the operations below and select the results to return.

  1. upload

    Send source code, tensors, bytes, compiled libraries, or files.

  2. get_function

    Select a named function or object from a module or library.

  3. run

    Call a function with arguments and keep its returned value.

  4. return

    Choose values, files, or folders for the response.

KCoral Engine

The execution backend for the protocol, from a single GPU host to a pool of compute nodes.

Unified entry pointRouterHealth & capacity
Compute node
KCoral serverWorker pool
CPU workGPU execution

SupervisorMonitor & restart server

Compute node
KCoral serverWorker pool
CPU workGPU execution

SupervisorMonitor & restart server

On a single host, clients can connect directly to the server. Add a router to share one address across nodes.

Coordinate GPU access

GPU stages run exclusively. Functions marked cpu_only=True release GPU access while they run.

Isolate request workspaces

With bubblewrap enabled, each request gets a private writable workspace and read-only runtime dependencies.

Reuse unchanged uploads

Content-addressed caches avoid sending the same code, data, and compiled libraries repeatedly.

Quick Start

Start the GPU host first, then install the client and run a program.

GPU Host

Linux · NVIDIA GPU and driver

Install

Set up the server environment.

System requirements
GPU host · terminal
python3 -m venv .venv
source .venv/bin/activate
python -m pip install 'kcoral[server]'

Start the server

Wait until it is ready.
Keep this terminal running.

Server setup
GPU host · terminal
kcoral server --device gpu --gpus 0 \
  --host 0.0.0.0 --port 8000

Client

No local GPU required

Install

Open a separate terminal for the client.

Client setup
Client · terminal
python3 -m venv .venv-client
source .venv-client/bin/activate
python -m pip install kcoral

Run a program

Save as example.py.
Run a calculation on the GPU host and return the result.

To connect from another machine, replace localhost with the GPU host's address.

Program guide
Client · example.py
from kcoral import Client, Program

program = Program()
module = program.upload(kind="module", source="""
def add_one():
    import torch
    return torch.arange(4, device="cuda") + 1
""")
fn = program.get_function(module=module, name="add_one")
output = program.run(fn=fn, args=[])
program.return_(key="output", value=output)

with Client("http://localhost:8000") as client:
    result = client.execute(program)
    print(result["output"])  # [1 2 3 4]

Adoptions