How the exercises work
Each exercise is a zip file you download, unpack, and run against your own GPU.
What you need
- An NVIDIA GPU and a CUDA toolkit, with
nvccon yourPATH. - Python 3.10 or newer. The runner uses only the standard library, so there is
nothing to
pip install. compute-sanitizerandncu(Nsight Compute) for thesanitizerandprofilecommands. Both ship with the CUDA toolkit.
No GPU? Use the submission toolkit
If you do not have an NVIDIA GPU in front of you, the gpu-submit toolkit
runs your CUDA on a datacentre GPU instead. You write the code on your laptop;
it is compiled and executed on a B200 or RTX PRO 6000 Blackwell box and the
results come back to your terminal. Python 3.11+ is the only local
requirement — no GPU, no CUDA toolkit, no nvcc, no Docker.
unzip gpu-submit.zip
cd gpu-submit
./setup.sh # paste the API key you were handed (gfaas_ws_...)
./runcu examples/device-query.cu
You need an API key to reach the GPU boxes — it is handed out with the course,
and setup.sh prompts for it. The bundled README.md covers submitting
exercises with ./submit, choosing a box with --gpu, and profiling remotely.
What is in the bundle
Unpacking a starter zip gives a self-contained directory:
axpy.cu the file you edit — a stub with the signature to implement
tester.cu the harness that allocates data, calls your code, and checks it
run.py entry point for every command
exercise.py the grading rules for this exercise
tests/ correctness cases, one text file each
benchmarks/ timing cases, one text file each
runner/ the runner itself; you never need to touch it
You edit only the exercise source (axpy.cu above). Both the kernel and its
launch are yours to write unless the exercise says otherwise; tester.cu calls
the host-side function whose signature the stub gives you.
A test case is a text file of parameters — array size, distributions, alignment offsets, seed, timeout — so you can read what a failing case did, and add cases of your own by copying one and editing it.
Running it
From inside the unpacked directory:
python run.py test
That compiles your file and runs the cases in tests/ in order, stopping at
the first failure. Passing cases are silent unless you ask for them with
python run.py -v test. The commands are:
| Command | What it does |
|---|---|
test |
Compile and check correctness against tests/ |
benchmark |
Time the cases in benchmarks/, optimized build |
sanitizer |
Re-run the tests under compute-sanitizer to catch invalid memory access |
profile |
Run ncu on the benchmark cases and summarize the result |
grade |
test, then sanitizer, then benchmark, stopping at the first failure |
compile |
Build the executable without running anything |
ptx / sass |
Emit the PTX or the disassembled SASS for your kernel |
Each command takes an optional list of cases, so python run.py test tests/05*
re-runs a single failing case. Options:
--file OTHER.cu— build a different source file, so several attempts can sit side by side.--arch 90— target a specific SM version instead of the detected one.--json PATH— also write the results as JSON.
python run.py with no command prints the same list.
What the runner tells you
A failing test names the case, its parameters, and the elements that came out wrong: index, inputs, the value your kernel produced, the value expected. Cases run from easy to awkward. A crash, a timeout, or a compile error is reported as such, with the compiler's or the program's own output.
benchmark reports the time per case. Where the exercise is about throughput,
it also reports how many MiB crossed the DRAM boundary, the bandwidth that
works out to, and what fraction of your device's peak that is.
profile prints duration, instruction count, DRAM and SM throughput, and
occupancy, and writes a full .ncu-rep report for the Nsight Compute UI. On
the benchmarks an exercise marks as its key cases, it adds feedback written for
that exercise: the kernel moved more memory than the problem requires, its
bandwidth is below what a straightforward solution reaches, or it has arrived
at the level the exercise aims for. The thresholds are per exercise and per
architecture.