Your First Program

Start a KCoral server, then submit a program that adds one to a four-element tensor on its GPU and returns the result.

Required hardware

You need one Linux machine with an NVIDIA GPU and a compatible driver. Follow Install the server and check the system requirements; the server installation also includes the client. The example uses PyTorch and does not compile a custom kernel.

The steps below run the server and client on that same machine, in two terminals. A CPU compilation server cannot run this program: the uploaded tensor requires GPU support, and the function also checks that it is on a GPU before doing arithmetic. Activate the same Python environment in both terminals.

Launch the server

Warning

KCoral allows clients to execute arbitrary code on its workers. Only allow trusted clients to access your KCoral server or Router. Deploy on a trusted, isolated network and never expose these endpoints to the public internet. Run workers in a sandbox with restricted permissions and access to host resources.

In the first terminal, start one worker on GPU 0:

kcoral server --device gpu --gpus 0 --workers-per-gpu 1 --host 127.0.0.1 --port 8000

GPU 0 is the first GPU listed by nvidia-smi. Wait for server startup to finish and leave this terminal running. The client will connect to http://127.0.0.1:8000.

Submit the program

In a second terminal, save the following complete program as first_program.py:

"""Upload a tensor, add one on the GPU, and return the result."""

import os

import numpy as np

from kcoral import Client, Program

SOURCE = """
def add_one(x):
    if not x.is_cuda:
        raise RuntimeError("This program requires a GPU tensor")
    return x + 1
"""


def build_program() -> Program:
    program = Program()
    module = program.upload(kind="module", source=SOURCE)
    add_one = program.get_function(module=module, name="add_one")
    x = program.upload(kind="tensor", value=np.arange(4, dtype=np.float32))
    y = program.run(fn=add_one, args=[x])
    program.return_(key="output", value=y)
    return program


def main() -> None:
    with Client(os.environ.get("KCORAL_URL", "http://127.0.0.1:8000")) as client:
        result = client.execute(build_program(), timeout_seconds=30)
    if not result.completed:
        raise SystemExit(f"Program failed: {result.error}")
    np.testing.assert_array_equal(result.results["output"], np.arange(1, 5, dtype=np.float32))
    print(result.status)
    print(result.results["output"])


if __name__ == "__main__":
    main()

You can also download first_program.py. Run it from the directory where you saved it:

KCORAL_URL=http://127.0.0.1:8000 python first_program.py

Expected output:

COMPLETED
[1. 2. 3. 4.]

The program checks that the request completed and that the returned array has the expected values. When you finish, press Ctrl+C in the server terminal to stop it.

Use a remote server

To submit from another machine, start the server with --host 0.0.0.0 so it listens beyond the local machine. Install the client on the submitting machine, save the same program there, and set KCORAL_URL to the server’s reachable address on port 8000. The client machine does not need a GPU. See Launch the server for more configuration.

How it works

The Program calls describe the work; client.execute() submits it. The server then executes the instructions in order:

  1. Load the uploaded source defining add_one.

  2. Select that function from the module.

  3. Transfer the uploaded NumPy input to the GPU.

  4. Call add_one with that tensor.

  5. Return the selected output. The client decodes it as a CPU NumPy array and checks the values.

The example also checks result.completed before reading the output. Instruction failures are returned as data in result.error; connection failures and request errors raise the exceptions documented in the Python API.

Continue to Write a client program for building programs and reading results, or Benchmark a Kernel with KCoral to compile a custom kernel, check correctness and measure its execution.