KCoral Protocol¶
KCoral exposes two client HTTP endpoints. A program is an ordered list of instructions, executed in one request with no persistent session handles. This page describes a direct server. The Router preserves the execution protocol while adding node selection and routing metadata.
Endpoints¶
POST /execute¶
Submit one program. The request body uses multipart/form-data with a JSON
program part and optional binary data parts. The request has no query fields.
See the complete example below.
Header |
Behavior |
|---|---|
|
Identifies each HTTP attempt and matches the result/error |
|
Identifies the node selected by the Router. Send it back as a cache-retry preference; a missing or unavailable preference falls back to another eligible node. |
Part |
Content type |
Required |
Meaning |
|---|---|---|---|
|
|
yes |
The program object below |
|
|
when not cached |
Raw tensor, byte, file or library content referenced by an upload |
The outer content type must include a multipart boundary. Each part needs one
Content-Disposition: form-data; name="..." header and one Content-Type
header. Send binary content as raw bytes, not base64; nested multipart bodies
are unsupported. Encode the program JSON as UTF-8. Part order is unrestricted.
SHA-256 is the content hash used to identify binary data. <sha256> is its
lowercase 64-character hexadecimal digest over the raw bytes.
Program field |
Type |
Required |
Meaning |
|---|---|---|---|
|
array |
yes |
Nonempty list of operations executed in order |
|
object |
no |
GPU count, execution timeout and captured output limits; see Options |
Each supplied binary part must be referenced by an upload; multiple uploads can share one part by naming the same digest. The server rejects duplicate or malformed part names, wrong content types and hashes that do not match the supplied bytes. JSON objects reject duplicate keys, and protocol objects reject unknown fields. Numeric options and JSON arguments reject non-finite numbers such as NaN and Infinity; argument objects may use arbitrary string keys.
Options
All options fields are optional. Timeout and output defaults and maxima are
server-configurable; values above the maxima are clamped. Booleans are not
accepted as numbers.
Field |
Type |
Default |
Meaning and limit |
|---|---|---|---|
|
integer |
omitted |
Use 1–8 GPUs with instruction-level leasing; omission uses the server’s ordinary GPU or CPU worker pool |
|
number |
|
Finite, strictly positive execution budget in seconds; maximum |
|
integer |
|
Non-negative capture limit per stdout/stderr stream; maximum |
The execution budget excludes waiting for a worker and waiting for GPU leases.
With explicit gpu_count, worker initialization and interpreter teardown count
against it. Admission timeouts and body-size limits are configured by the
server, not by request
options. Body-size limits include multipart framing.
See Response for HTTP statuses, result fields, and value encodings.
GET /health¶
Read endpoint health, request load, and the compilation environment.
GET /health HTTP/1.1
Host: 127.0.0.1:8000
Example response from a GPU server with two workers:
{
"status": "ok",
"instance_id": "09dc4eaa-a8b1-46cf-b5fb-a3448dcd7ca6",
"started_at": "2026-08-29T18:42:11.019012Z",
"gpu_count": 1,
"load": {
"request_capacity": 2,
"requests_in_progress": 2,
"requests_waiting": 3
},
"target": {"arch": "sm_100a"},
"versions": {"torch": "2.14.0+cu132", "cuda": "13.2", "tvm_ffi": "0.1.13.post2"}
}
Version strings above are illustrative; use the values returned by your server.
The GPU runtime reports torch, cuda (the CUDA version used to build PyTorch),
and available tvm, tvm_ffi, triton, cutlass and flashinfer versions.
Optional entries can be absent; this is not an inventory of every installed
compiler or command-line tool. CPU servers report the available tvm_ffi version.
Field |
Type |
Meaning |
|---|---|---|
|
string |
Direct server: |
|
string |
Changes on each endpoint restart |
|
string |
Endpoint startup time in UTC (RFC 3339) |
|
integer or null |
Configured GPUs; |
|
integer |
Serviceable request capacity, occupied and free |
|
integer |
Assigned requests, including compilation, GPU waiting, and cleanup |
|
integer |
Requests awaiting assignment at this endpoint |
|
object |
Compilation target, including |
|
object |
Runtime and toolchain version strings |
Direct-server capacity counts workers, excluding background replacements, and becomes zero during shutdown. Router capacity counts execution connections on eligible nodes; its request counts cover submissions through that router.
Counts can change before submission. During recovery or cleanup, in-progress
requests may exceed capacity. Assigned requests can wait for GPU access even
when requests_waiting is zero.
On the router, instance_id and started_at describe the router process.
For choosing a compilation target, see Remote Compilation.
Operations¶
Every instruction is a JSON object with an op field. The four values are
upload, get_function, run and return.
Common field |
Meaning |
|---|---|
|
Required operation name; determines the accepted fields |
|
Required, nonempty, unique identifier when the operation produces a handle; absent for |
|
A reference value naming an earlier handle; it is not a top-level instruction field |
A handle names a value inside this request. upload, get_function and run
produce handles. References must have exactly the $ref
key and cannot refer forward or cross request boundaries. Each operation below
specifies where references are resolved. Return keys are unique in a separate
namespace from instruction identifiers. Each field table is exhaustive:
unlisted fields are rejected.
upload¶
Upload source or binary data and produce a handle.
{
"op": "upload",
"id": "module",
"kind": "module",
"source": "def add_one(x): return x + 1"
}
Each field is required for the listed kinds.
Field |
Kinds |
Notes |
|---|---|---|
|
all |
|
|
all |
Unique handle name |
|
all |
|
|
module |
Python source executed to define a module |
|
tensor, bytes, file, library |
SHA-256 of the raw bytes |
|
file |
Relative destination in the request working directory |
|
tensor |
Tensor data type |
|
tensor |
Tensor shape |
Kind |
Content and result |
|---|---|
Inline Python |
|
Raw contiguous row-major bytes, copied to the assigned GPU. The byte length must equal |
|
The blob’s bytes unchanged in CPU memory. Pass them to uploaded Python to parse files or other binary formats. |
|
|
Materialize the blob at |
|
A precompiled TVM FFI shared library for the server’s platform, loaded with |
Tensor dtype must be one of bool, uint8, int8, int16, int32,
int64, float16, float32, float64, bfloat16, float8_e4m3fn, or
float8_e5m2. Elements use little-endian byte order. shape is an array of
non-negative integers, excluding booleans: [] describes a scalar, and a zero
dimension describes an empty tensor. Uploads use the worker’s current CUDA
device, initially logical device 0; there is no upload field selecting a GPU.
CPU workers accept module, bytes and file uploads. Tensor and library uploads
fail at execution with unavailable on a CPU worker.
For each uncached binary upload, supply the bytes in a multipart part named
blob:<sha256>. For example:
[
{"op": "upload", "id": "input", "kind": "tensor", "blob": "<sha256>",
"dtype": "float16", "shape": [32, 128]},
{"op": "upload", "id": "raw_bytes", "kind": "bytes", "blob": "<sha256>"},
{"op": "upload", "id": "file_path", "kind": "file", "blob": "<sha256>", "path": "data/tensor.bin"}
]
File upload rules
Rule |
Behavior |
|---|---|
Path |
Relative POSIX path naming a file. Rejects |
Normalization |
Removes |
Length |
Each component is at most 255 UTF-8 bytes; the normalized path is at most 4096 bytes. Empty paths and the workspace root ( |
Conflicts |
Normalized paths must be unique and cannot conflict as a file and directory; a program cannot upload both |
Filesystem access |
Creates a regular file with mode |
Workspace |
A fresh temporary working directory per request, owned by the parent process and removed after completion, failure, timeout, or worker crash. |
Lifetime |
Blob cache entries remain available after materialized files are removed. |
An upload cannot overwrite a file already created by an earlier instruction.
With bubblewrap isolation, .kcoral is reserved for runtime files and uploads
there fail with runtime. Materialization failures are instruction errors,
whereas invalid literal paths and declared upload conflicts are rejected before
execution with HTTP 400.
Library¶
A library is an already-built TVM FFI module, whether its bytes came from the client’s toolchain or a preceding CPU-server request. Uploading it requires no compilation:
{
"op": "upload",
"id": "kernels",
"kind": "library",
"blob": "<sha256>"
}
blob names the bytes of an ELF shared object for the server’s platform. The
server loads it with tvm_ffi.load_module and binds the resulting module.
Later get_function instructions may bind any number of its TVM FFI exports. A library
that cannot be loaded, or a requested function that is absent, fails with a
compile error. Nothing else about the object is inspected, so any producer
TVM FFI can load is accepted. Three are usual:
Producer |
Export |
Server requirement |
|---|---|---|
C++ / |
|
TVM FFI |
|
Embedded module blob |
TVM installed, including the loader registered by its CUDA runtime |
CuTeDSL with |
|
A compatible CuTeDSL runtime. Check |
The exported function determines its argument types. Build device code for
GET /health’s target architecture and provide compatible runtime dependencies.
KCoral does not validate the library’s target architecture before loading it;
an incompatible device image can fail at launch.
For build examples, see local compilation and CPU-server compilation.
get_function¶
Select a named object from an earlier module or library upload.
{
"op": "get_function",
"id": "add_one",
"module": {"$ref": "module"},
"name": "add_one"
}
Field |
Type |
Required |
Notes |
|---|---|---|---|
|
string |
yes |
|
|
string |
yes |
Handle for the callable |
|
|
yes |
Earlier |
|
string |
yes |
Non-empty function or object name |
|
boolean |
no |
Defaults to |
Source |
Selection behavior |
|---|---|
Python module |
Looks up |
TVM-FFI library |
Calls the loaded module’s |
Module and function handles are request-local capabilities and cannot be returned in a response.
A function declared cpu_only touches no GPU. A run of that handle releases
the worker’s GPU lease first, and a CUDA runtime or driver API call from it, as
seen by CUPTI, fails the instruction with error kind gpu_access. The check is
best effort: it sees a call only after it has begun, and none from a child
process. The declaration applies to the handle as a run target only.
On GPU workers, module uploads execute their imports and top-level Python under
the GPU lease, regardless of how functions are later selected. Tensor and library
uploads, get_function, ordinary run calls and value returns also acquire the
lease. File uploads and file/folder returns release it after synchronizing;
byte uploads leave lease ownership unchanged. Final runtime cleanup reacquires
the lease even after a CPU-only tail.
run¶
Call a function and bind the value it returns. Here add_one and x refer to
earlier instructions.
{
"op": "run",
"id": "y",
"fn": {"$ref": "add_one"},
"args": [{"$ref": "x"}]
}
Field |
Type |
Required |
Meaning |
|---|---|---|---|
|
string |
yes |
|
|
string |
yes |
Unique handle for the result |
|
reference |
yes |
An earlier callable handle, |
|
array |
no |
Positional arguments, default |
Value in |
Passed to the callable |
|---|---|
Top-level |
Value of the earlier handle |
Reference-shaped object nested in a list or object |
Literal JSON; not resolved |
Other JSON value |
Literal value |
Selecting an object with get_function does not prove it is callable; using a
non-callable object here fails at execution. run binds the computed value but
does not include it in the response; add a return to expose it.
Uploaded Python can perform tasks such as allocation, compilation, correctness
checks and measurement. A callable returned by a run can be used by a later
run.
return¶
Select an earlier value for the response.
{
"op": "return",
"key": "output",
"value": {"$ref": "y"}
}
Field |
Type |
Required |
Notes |
|---|---|---|---|
|
string |
yes |
|
|
string |
yes |
Nonempty, unique key in the response |
|
|
For value returns |
Earlier handle to return; omit |
|
string |
For file/folder returns |
|
|
string or |
For file/folder returns |
Workspace path, supplied literally or through an earlier handle |
return snapshots its value at that instruction and does not stop execution.
It creates no handle. See Errors for partial-result behavior.
File and folder returns
{"op": "return", "key": "report", "kind": "file", "path": "outputs/report.txt"}
To return a folder whose path is held in an earlier register:
{"op": "return", "key": "debug", "kind": "folder", "path": {"$ref": "output_path"}}
Rule |
Behavior |
|---|---|
Paths |
Follow file-upload rules, relative to the original request workspace even if code changes cwd. |
Snapshot |
Contents are captured at that instruction, without holding the GPU lease. Folders include hidden files and empty directories; original metadata is omitted. |
Rejected content |
Symlinks, special files, repeated directories, and observable changes during reads |
Failures |
Missing paths, wrong types, invalid runtime paths, read failures, and collection limits fail that return with |
Size limit |
Contents are buffered and count against the server response-size limit. |
Multi-GPU execution¶
Set options.gpu_count to an integer from 1 through 8 to run the program
once on that many GPUs. CUDA_VISIBLE_DEVICES exposes the assigned set as
logical devices 0 through gpu_count - 1; the response’s gpu_ids reports
physical devices. The script owns process creation, communication, and
synchronization. KCoral does not broadcast instructions or create a communication
group. Registers belong to the request’s interpreter.
Single- and multi-GPU programs use the same worker and instruction-level leasing.
Before a GPU instruction, the worker acquires its complete device set atomically.
Before cpu_only execution it synchronizes all devices and releases the set;
a later GPU instruction reacquires the same devices. Waiting for devices
follows the timeout rules in Options. Releasing a lease does not
move or discard this program’s tensors.
GPU subprocesses must finish before their run returns. Do not mark a GPU
launcher CPU-only.
On timeout or crash, the worker cleans up its process tree before abandoning its leases. Unverified cleanup makes the affected devices unavailable and stops the server pool from accepting new requests. Leftover descendants are terminated and fail the request. Shutdown cancels queued allocations and drains running requests.
Explicit gpu_count requests use a fresh worker with the assigned set;
omitting the option uses the configured single-GPU workers. This option requires
Linux and a direct connection to a GPU server; the router does not select by GPU
count. Counts exceeding server capacity, or requests to a CPU server, fail with
HTTP 400. See multi-GPU examples.
Upload Caching¶
Tensor, byte and library uploads use the memory cache; file uploads use the persistent file cache. Both are keyed by the raw content’s SHA-256. Module source travels inline and is not cached. Caches retain uploaded bytes, not execution state. Resolved bytes remain usable by an admitted request even if evicted.
Cache capacity, eviction and storage settings are described in server cache configuration.
To use a cached upload, send its hash but omit its binary part. If required bytes are absent, the server returns before executing any instruction:
{
"status": "CACHE_MISS",
"request_id": "7f61b94e-034a-4e80-b67d-eca52bb952cc",
"missing_blobs": ["<sha256>"]
}
Resend the same program with the listed parts. The Python client does this automatically, then makes one final attempt with every local blob if another cache miss occurs. Cache retention is an optimization rather than a guarantee.
The same digest may exist in either or both cache categories. A request using that digest for both a file and a tensor, byte string or library can share the resolved bytes; newly supplied content is offered to each referenced category. A hit in one category does not generally guarantee a hit in the other.
Response¶
HTTP |
Body |
Meaning |
|---|---|---|
200 |
|
Program completed |
200 |
|
Instruction or cleanup failure; see Errors |
200 |
|
Referenced blobs are missing; program did not run |
400 |
|
Malformed request or program, including duplicate JSON keys and NaN/Infinity |
413 |
|
Request body exceeds the server’s size limit |
503 |
|
No worker is available; includes |
504 |
|
Execution timed out |
500 |
|
Worker failure outside an instruction, or server failure |
500 |
|
Results exceed the server’s response-size limit |
A 200 response carries the fields below. A non-200 carries the smaller error body described under Errors instead.
Field |
Type |
Present |
Notes |
|---|---|---|---|
|
string |
always |
|
|
string |
always |
Also sent as |
|
number |
run |
Worker wait time |
|
number |
run |
Time after worker assignment, including GPU waits, execution, result encoding and cleanup |
|
number |
run |
Waiting for the assigned GPU set |
|
number |
run |
Wall-clock time holding the GPU set, not summed GPU activity time |
|
array of integers |
run with explicit |
Allocated physical devices, in logical-device order |
|
integer |
run with explicit |
Number of allocated devices |
|
object |
run |
Captured returns, subject to failure behavior |
|
object |
|
See Errors |
|
array |
|
Blob hashes the server does not hold |
|
string |
run |
Captured standard output |
|
string |
run |
Captured standard error |
|
boolean |
run |
Whether captured stdout exceeded |
|
boolean |
run |
Whether captured stderr exceeded |
“run” marks fields present whenever the worker returned an outcome, so on both
COMPLETED and FAILED but not on CACHE_MISS.
Captured streams are decoded as UTF-8 with replacement for invalid bytes. The capture limit applies to raw bytes before decoding. With capture disabled, both strings are empty and both truncation flags are false.
For a successful program, HTTP status is 200:
{
"status": "COMPLETED",
"request_id": "7f61b94e-034a-4e80-b67d-eca52bb952cc",
"queue_ms": 0.4,
"elapsed_ms": 812.6,
"lease_wait_ms": 0.0,
"lease_held_ms": 1.4,
"results": {
"timing": {
"type": "object",
"value": {
"latency_ms_median": {"type": "number", "value": 0.0073}
}
}
},
"stdout": "",
"stderr": "",
"stdout_truncated": false,
"stderr_truncated": false
}
elapsed_ms includes lease_wait_ms and lease_held_ms; the remainder is
time without a GPU lease. Holding a lease reserves the GPU and includes host
work during that period. Kernel latency is measured separately by the program.
Value encoding¶
Type |
Encoding |
|---|---|
null |
|
boolean |
|
integer |
|
number |
|
string |
|
array |
|
object |
|
bytes |
|
file |
|
folder |
|
tensor |
|
Arrays and objects recursively contain encoded values. Object keys are unique
strings with no ordering semantics. Numbers must be finite. Python lists and
tuples both encode as array.
If no bytes, tensor, or file appears, the response is application/json. Otherwise
it is multipart/form-data:
Part |
Content type |
Required |
Notes |
|---|---|---|---|
|
|
yes |
Response metadata and value tree |
|
|
conditional |
Raw bytes for a bytes, tensor, or file node |
Binary parts use depth-first numbering. Clients use part to locate data and
verify sha256. Tensor data is C-contiguous, row-major, and little-endian; its
length must match dtype and shape. A file’s size is a non-negative integer
(not a boolean) and must match its binary part length.
Folder files maps sorted relative paths to file values; directories lists
sorted directory paths, including every ancestor. The selected root is implicit.
Paths must be canonical under file-upload rules, unique, and free of file/directory
conflicts. An empty folder has empty files and directories. Each file uses
its own binary part with the usual depth-first numbering and integrity checks.
Errors¶
An ordinary instruction failure stops the program and preserves earlier returns. A worker crash loses those returns and captured streams; HTTP-level failures carry no partial results.
The error object in a FAILED response has this form:
{
"kind": "correctness",
"message": "outputs differ: max_abs_err=0.5 exceeds atol=0.001",
"instruction_index": 6,
"instruction_op": "run",
"instruction_id": "check",
"traceback": "Traceback (most recent call last):\n ..."
}
The failing instruction itself contributes nothing: a return that fails while
encoding adds neither a results entry nor binary parts.
Field |
Type |
Notes |
|---|---|---|
|
string |
See kinds below |
|
string |
Human-readable description |
|
integer |
Zero-based position in |
|
string |
|
|
string | null |
The instruction’s |
|
string |
Server-side traceback, retaining at most the final 8192 characters; can be empty for native crashes or cleanup errors |
Instruction error kinds are parse, compile, runtime, gpu_access,
correctness, serialization, unavailable, and engine. An unhandled Python
exception from a run, including AssertionError, is reported as engine;
correctness is used when the harness explicitly raises that execution error.
A gpu_access error means a cpu_only function entered the CUDA API; it adds
cuda_call, location, and interfered_request_id, and its traceback is
the stack at that call. interfered_request_id identifies the other request
holding the worker’s GPU set at detection time, or is null when there is no
single identifiable holder.
ERROR is not a program outcome, so its body is much smaller: status,
request_id, and an error of kind and message only, with no results,
timings, or captured output.
{
"status": "ERROR",
"request_id": "7f61b94e-034a-4e80-b67d-eca52bb952cc",
"error": {"kind": "busy", "message": "server saturated"}
}
Its kind is parse, request_too_large, busy, timeout, engine, or
response_too_large — a separate set from the instruction kinds above.
Failure |
Outcome and recovery |
|---|---|
Native uploaded code terminates its worker |
|
CUDA error poisons the worker’s context |
|
Non-sticky CUDA launch error found while draining the request |
|
Timeout, worker failure outside an instruction, or server failure |
|
Cleanup can report a failure after the final instruction, attributed to the last instruction observed. Worker replacement also follows the configured request limit, which defaults to one request per process.
Router errors¶
The Router forwards server outcomes and can also return its own ERROR body
with the same status, request_id and error fields:
HTTP |
|
Meaning |
|---|---|---|
400 |
|
Invalid request headers or a failure reading the client’s body |
413 |
|
Router request-size limit exceeded |
503 |
|
Router queue full; includes |
503 |
|
Wait for compatible node capacity expired; includes |
502 |
|
Node connection failed or returned an invalid response; execution outcome may be unknown |
For retry and disconnect behavior, see Router recovery.
Example¶
{
"instructions": [
{
"op": "upload",
"id": "harness",
"kind": "module",
"source": "def main(x): return x + 1"
},
{
"op": "get_function",
"id": "main",
"module": {
"$ref": "harness"
},
"name": "main",
"cpu_only": true
},
{
"op": "run",
"id": "output",
"fn": {
"$ref": "main"
},
"args": [
41
]
},
{
"op": "return",
"key": "output",
"value": {
"$ref": "output"
}
}
],
"options": {
"timeout_seconds": 30
}
}
This program uploads its own harness and returns 42; it needs no binary parts.
The same upload/get_function/run shape supports GPU harnesses and compiler tools.
The Python package constructs request parts, hashes and response values for you.
Task |
Python client reference |
|---|---|
Submit a first request |
|
Construct programs and manage clients |
|
Upload files and folders |
|
Receive files and folders |
Results decode to |