How the exercises work

Each exercise is a zip file you download, unpack, and run against your own GPU.

What you need

No GPU? Use the submission toolkit

If you do not have an NVIDIA GPU in front of you, the gpu-submit toolkit runs your CUDA on a datacentre GPU instead. You write the code on your laptop; it is compiled and executed on a B200 or RTX PRO 6000 Blackwell box and the results come back to your terminal. Python 3.11+ is the only local requirement — no GPU, no CUDA toolkit, no nvcc, no Docker.

Download gpu-submit.zip

unzip gpu-submit.zip
cd gpu-submit
./setup.sh          # paste the API key you were handed (gfaas_ws_...)
./runcu examples/device-query.cu

You need an API key to reach the GPU boxes — it is handed out with the course, and setup.sh prompts for it. The bundled README.md covers submitting exercises with ./submit, choosing a box with --gpu, and profiling remotely.

What is in the bundle

Unpacking a starter zip gives a self-contained directory:

axpy.cu          the file you edit — a stub with the signature to implement
tester.cu        the harness that allocates data, calls your code, and checks it
run.py           entry point for every command
exercise.py      the grading rules for this exercise
tests/           correctness cases, one text file each
benchmarks/      timing cases, one text file each
runner/          the runner itself; you never need to touch it

You edit only the exercise source (axpy.cu above). Both the kernel and its launch are yours to write unless the exercise says otherwise; tester.cu calls the host-side function whose signature the stub gives you.

A test case is a text file of parameters — array size, distributions, alignment offsets, seed, timeout — so you can read what a failing case did, and add cases of your own by copying one and editing it.

Running it

From inside the unpacked directory:

python run.py test

That compiles your file and runs the cases in tests/ in order, stopping at the first failure. Passing cases are silent unless you ask for them with python run.py -v test. The commands are:

Command What it does
test Compile and check correctness against tests/
benchmark Time the cases in benchmarks/, optimized build
sanitizer Re-run the tests under compute-sanitizer to catch invalid memory access
profile Run ncu on the benchmark cases and summarize the result
grade test, then sanitizer, then benchmark, stopping at the first failure
compile Build the executable without running anything
ptx / sass Emit the PTX or the disassembled SASS for your kernel

Each command takes an optional list of cases, so python run.py test tests/05* re-runs a single failing case. Options:

python run.py with no command prints the same list.

What the runner tells you

A failing test names the case, its parameters, and the elements that came out wrong: index, inputs, the value your kernel produced, the value expected. Cases run from easy to awkward. A crash, a timeout, or a compile error is reported as such, with the compiler's or the program's own output.

benchmark reports the time per case. Where the exercise is about throughput, it also reports how many MiB crossed the DRAM boundary, the bandwidth that works out to, and what fraction of your device's peak that is.

profile prints duration, instruction count, DRAM and SM throughput, and occupancy, and writes a full .ncu-rep report for the Nsight Compute UI. On the benchmarks an exercise marks as its key cases, it adds feedback written for that exercise: the kernel moved more memory than the problem requires, its bandwidth is below what a straightforward solution reaches, or it has arrived at the level the exercise aims for. The thresholds are per exercise and per architecture.

The exercises

Exercise Part
1.7 Copying a rectangle Part 1
2.7 SAXPY Part 2
3.5 Embedding bags Part 3
4.4 Matrix multiplication Part 4
5.4 Parallel Reduction Part 5
5.7 Matrix Transpose Part 5
6.7a AXPY with bfloat16 Part 6
6.7b AXPY with misaligned data Part 6
7.4 Row-wise Softmax in BF16 Part 7
8.5 Max pooling Part 8