Benchmark a Kernel with KCoral¶
Evaluating a kernel means answering two questions: does it produce the right output, and how long does it take to run? In this tutorial, you will send a small add-one kernel to a KCoral GPU server, check its output against a reference, and retrieve a timing report. The client describes the experiment; compilation, correctness checks, and measurement happen on the server.
The example uses TVM to compile a TIRx kernel on the server’s CPU, then
KCoral’s benchmark() helper to measure it on the GPU.
You can substitute your own compiler or measurement code and use the same
program structure for other kernel languages.
Prerequisites¶
On the server machine, install the
GPU worker environment
and launch the server. This example
uses TIRx, TVM’s Python-embedded kernel language, so the server needs TVM and
the CUDA toolkit’s nvcc compiler for CPU compilation, PyTorch for tensors and
correctness checks, and CUPTI for measurement.
On the client machine, install the KCoral client, which includes NumPy for preparing the input array. This example needs no local GPU, CUDA toolkit, or TVM installation: the client sends source and data to the server.
Run a complete example¶
Start with examples/benchmark_kernel/benchmark_kernel.py, which adds one to each of 256
float32 values. From the repository checkout, run it against your GPU server:
KCORAL_URL=http://127.0.0.1:8000 python examples/benchmark_kernel/benchmark_kernel.py
"""Compile on a remote server's CPU, then check and time the kernel on its GPU."""
import os
import numpy as np
from kcoral import Client, Program
SOURCE = r"""
from __future__ import annotations
import os
from pathlib import Path
import torch
import tvm
import tvm_ffi
from tvm.script import tirx as T
from kcoral.builtins import benchmark
@T.jit
def add_one(A: T.Buffer((N,), "float32"), B: T.Buffer((N,), "float32"), *, N: T.constexpr):
T.device_entry()
i = T.cta_id([N])
Show less: benchmark_kernel/benchmark_kernel.py
"""Compile on a remote server's CPU, then check and time the kernel on its GPU."""
import os
import numpy as np
from kcoral import Client, Program
SOURCE = r"""
from __future__ import annotations
import os
from pathlib import Path
import torch
import tvm
import tvm_ffi
from tvm.script import tirx as T
from kcoral.builtins import benchmark
@T.jit
def add_one(A: T.Buffer((N,), "float32"), B: T.Buffer((N,), "float32"), *, N: T.constexpr):
T.device_entry()
i = T.cta_id([N])
t = T.thread_id([1])
B[i] = A[i] + 1.0
def compile_kernel(arch):
# An explicit architecture avoids querying the GPU during compilation.
target = tvm.target.Target({"kind": "cuda", "arch": arch})
module = tvm.IRModule({"add_one": add_one.specialize(N=256)})
# NVRTC can initialize CUDA; use the CPU-only nvcc subprocess instead.
key = "TVM_CUDA_COMPILE_MODE"
previous = os.environ.get(key)
os.environ[key] = "nvcc"
try:
with target:
compiled = tvm.compile(module, target=target, tir_pipeline="tirx")
compiled.export_library("add_one.so")
finally:
if previous is None:
os.environ.pop(key, None)
else:
os.environ[key] = previous
# Keep the executable alive: destroying its CUDA module can call CUDA too.
return compiled, "add_one.so"
def evaluate(compilation, src):
_executable, library = compilation
module = tvm_ffi.load_module(Path(library).resolve())
compiled = module["add_one"]
dst = torch.empty_like(src)
compiled(src, dst)
torch.testing.assert_close(dst, src + 1.0, rtol=1e-2, atol=1e-3)
return {"check": {"passed": True}, "timing": benchmark(compiled, src, dst)}
"""
def build_program(arch: str) -> Program:
program = Program()
module = program.upload(kind="module", source=SOURCE)
compile_kernel = program.get_function(module=module, name="compile_kernel", cpu_only=True)
evaluate = program.get_function(module=module, name="evaluate")
compilation = program.run(fn=compile_kernel, args=[arch])
values = np.arange(256, dtype=np.float32)
src = program.upload(kind="tensor", value=values)
report = program.run(fn=evaluate, args=[compilation, src])
program.return_(key="report", value=report)
return program
def main() -> None:
with Client(os.environ.get("KCORAL_URL", "http://127.0.0.1:8000")) as client:
result = client.execute(build_program(client.target()["arch"]), timeout_seconds=120)
if not result.completed:
raise SystemExit(f"Benchmark failed: {result.error}")
print(result.results["report"]["check"])
print(result.results["report"]["timing"])
if __name__ == "__main__":
main()
The client reads the GPU architecture from Client.target() and uploads the
kernel with separate compile_kernel and evaluate functions. It selects
compile_kernel with cpu_only=True, so compilation releases the GPU for
other requests. It then uploads the input tensor and calls evaluate with
the compilation result. This call keeps GPU access, loads the library, allocates an output tensor,
and checks its result against src + 1. It only benchmarks the kernel after
that check passes. On success, the script prints
{'passed': True} followed by the timing report; otherwise it reports the
program error. The successful output has this shape, with the timing dictionary
abridged here and the latency depending on your GPU:
{'passed': True}
{'latency_ms_median': <latency>, ...}
Inside compile_kernel, add_one.specialize(N=256) supplies the kernel’s
compile-time size. tvm.compile uses an explicit CUDA target architecture and
the nvcc subprocess backend, so compilation does not initialize CUDA in the
worker. It exports the compiled library as
add_one.so in the request’s workspace and returns the executable together
with its path. Keeping the executable alive defers CUDA module cleanup until
GPU access is available. The later GPU call loads that library without
repeating compilation.
Inside evaluate, benchmark(compiled, src, dst) measures GPU activity with
CUPTI. Keeping compilation and evaluation in separate calls lets the server
release GPU access for compilation and reacquire it for execution.
The rest of this tutorial looks at the choices behind that example: where to compile, how to overlap CPU and GPU work, and how to check and measure the kernel.
Choose where to compile¶
The first example compiles on the GPU server, keeping the experiment in one request. You can also compile elsewhere when you want to reuse an existing build environment or scale compilation separately from GPU execution.
On a GPU server¶
Compiling on the GPU server keeps the client lightweight and uses the server’s
installed toolchain. The example above compiles and executes in one request,
with only the compilation call marked cpu_only=True.
Other compiler APIs can initialize CUDA or load GPU modules, including
compile_tirx(). Calls that do this must keep GPU access
and leave cpu_only at its default of False. CUDA C, CuTeDSL and Triton can
use their own compiler APIs with the same distinction between CPU compilation
and GPU work. Client.health() reports installed versions.
On the client¶
If you already have the compiler on your client machine, you can keep the build
local and use the remote GPU for evaluation. The
examples/cpu_compile/library_upload.py example does this for a
CUDA C add-one kernel. It needs the CUDA toolkit, a host C++ compiler, and TVM FFI
with its C++ build dependencies on the client. See
local compilation setup.
Its build_library function writes the CUDA source to a temporary directory,
compiles it for the supplied GPU architecture, and reads the resulting shared
library as bytes:
def build_library(arch: str, directory: str) -> bytes:
"""Compile SOURCE for `arch` — the server's, not this machine's."""
import tvm_ffi.cpp
source_path = os.path.join(directory, "add_one.cu")
pathlib.Path(source_path).write_text(SOURCE)
# Whatever the local toolchain accepts belongs here; this freedom is the
# reason to upload a library rather than let the server build one.
library = tvm_ffi.cpp.build(
"add_one",
cuda_files=source_path,
extra_cuda_cflags=[
f"-gencode=arch=compute_{arch.removeprefix('sm_')},code={arch}",
"-O3",
],
build_directory=directory,
output=os.path.join(directory, "add_one.so"),
)
return pathlib.Path(library).read_bytes()
Here, SOURCE is the example’s CUDA code, including a TVM FFI export for
add_one. The client reads the remote GPU’s architecture before building:
arch = client.target()["arch"]
with tempfile.TemporaryDirectory() as directory:
library = build_library(arch, directory)
The program then uploads the compiled bytes and selects the exported function:
module = program.upload(kind="library", value=library)
kernel = program.get_function(module=module, name="add_one")
Later instructions call kernel with the input and output tensors, check the
result, and measure it. Compilation has already finished on the client, so the
server only needs to load and execute the library. See the
library protocol for export and linking
requirements. Run the complete client from the repository checkout:
KCORAL_URL=http://127.0.0.1:8000 python examples/cpu_compile/library_upload.py
On a CPU server¶
If a CPU server is available and the build does not require a GPU, you can use
it for compilation. It executes uploaded Python compilation code and returns
library bytes for a subsequent request to the GPU server. Supply the target architecture from
Client.target() on the GPU server. Upload CUDA source with upload_file and pass
its path and export names to your compilation function.
Remote Compilation walks through building CUDA C on a CPU server, then uploading the resulting library to a GPU server for execution and measurement.
Overlap CPU work with GPU execution¶
Some parts of an experiment, such as preparing source files or running a CPU-only compiler, do not need the GPU. While one request does that work, another could use the GPU to run a kernel.
KCoral coordinates this by granting one worker exclusive GPU access at a time,
called a GPU lease. By default, an uploaded function keeps that access for
its entire call. If the function does no GPU work, select it with
get_function(..., cpu_only=True). Before running it, the worker waits for its
outstanding GPU work to finish and releases its exclusive access, allowing
another request to use the GPU. Detected CUDA calls in the CPU-only function
fail with gpu_access; this is a best-effort check.
The cpu_only=True option applies only when run invokes the selected
function. Uploading a Python module is a separate step: the server executes
its imports and other statements outside function bodies to create the module.
For example, a tensor allocation written outside a function runs during upload.
That code may initialize CUDA or use the GPU, so KCoral gives the module upload
exclusive GPU access even if you later select one of its functions as CPU-only.
To let compilation overlap with another request’s GPU work, put the build in
a function that does not access the GPU and select it with cpu_only=True.
For example, it could run nvcc to produce a shared library and return the
library’s file path. A later function can take that path, load the library,
and launch the kernel. Leave cpu_only at its default of False for this
second function so KCoral gives it exclusive GPU access. Keeping the two steps
separate lets other requests use the GPU while the compiler runs.
Check correctness¶
Before trusting a timing result, check that the kernel computes the expected
answer. The first example uses torch.testing.assert_close to compare its
output against a reference; your uploaded code can do the same with tolerances
appropriate to your workload. You can also upload your own correctness-checking
functions or scripts to validate outputs against a reference or test properties
specific to your kernel. Return any reports you want to inspect. An assertion
failure stops subsequent instructions; results already selected by return_
remain available.
Measure GPU activity¶
Once correctness passes, measure how long the kernel’s GPU work takes.
kcoral.builtins.benchmark measures each call from its first GPU activity to
its last, including kernels, copies and memsets. Host work before and after those endpoints
is excluded; gaps between GPU activities are included. This supports functions
that launch multiple GPU operations.
To adjust the measurement, pass an optional configuration dict after the
callable’s arguments in your uploaded Python. Here, kernel is the callable
to measure, and src and dst are its input and output tensors:
from kcoral.builtins import benchmark
timing = benchmark(kernel, src, dst, {"warmup_ms": 25, "repeat_ms": 100, "flush_l2": True})
Those are the defaults. The time budgets determine iteration counts from an
initial estimate. Explicit warmup and repeat counts override their respective
budgets. L2 flushing happens before each call and outside its measured span.
The report contains latency_ms_median, latency_ms_mean, latency_ms_min,
latency_ms_max, warmup, repeat, flush_l2 and activities_stable. A false
activities_stable means calls did not all launch the same GPU activities.
The helper requires PyTorch and cupti-python in the worker environment.
Use this report to assess kernel latency. The request timing fields on
ProgramResult include other costs, such as waiting and compilation; see
reading results when
investigating a slow request.
For profiling with NCU or run-iket, or checking GPU memory access with Compute Sanitizer, use the builtin CLI tools. They run the tool on the server and retrieve its reports for you.