High-Performance Infrastructure

Kaiser I

EU & Land Niedersachsen Funding

Kaiser I is our scalable, high-performance GPU computing cluster designed to advance generative AI models, cognitive robotics, and circular systems engineering.

5

Total Nodes

1 Login, 4 Workers

16

NVIDIA GPUs

8x H200, 8x RTX 6000

€400k+

Funding

EU & Niedersachsen

Slurm

Scheduler

Resource manager

Project Description

The Kaiser I GPU Cluster

The Kaiser I cluster is established as the central digital platform at the Technical University of Clausthal to unlock the full potential of Generative Artificial Intelligence (GenAI) for systems engineering. The goal is to both advance fundamental AI methods and drive concrete engineering applications - ranging from resource-efficient design and circular production to cognitive robotics and sustainable mobility.

Led by Prof. Dr. Christian Bartelt at the Institute for Software and Systems Engineering (ISSE), Kaiser I provides a scalable, high-performance environment that strengthens research, teaching, and technology transfer while fostering regional and international innovation partnerships.

Inauguration of Kaiser I
Prof. Dr. Christian Bartelt and researchers powering up the new Kaiser I GPU Cluster at TU Clausthal.

Kaiser I funded by the European Union & Land Niedersachsen (EFRE/ESF/ELER)

Located at Rechenzentrum der TU Clausthal, Erzstraße 18

Led by Prof. Dr. Christian Bartelt (ISSE)

Open workshops and innovation hub access for industry partners

Hardware & Nodes

Cluster nodes and hardware layout

The cluster is structured into login and worker nodes containing specific GPU architectures.

2 Nodes · H200 GPUs

High-Performance Nodes

Two high-performance worker nodes, each hosting up to 8 NVIDIA H200 NVL GPUs (minimum 8 H200 GPUs total) interconnected via NVLink for maximum data throughput.

  • Up to 8 H200 NVL per node
  • Optimized for model fine-tuning up to 70B
  • Inference of Llama 3, Qwen 3, Gemma 3
2 Nodes · RTX 6000 GPUs

Mid-Performance Nodes

Two mid-performance worker nodes hosting up to 8 NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs. Tailored for resource-efficient training and vision models.

  • Blackwell Max-Q GPU architecture
  • Optimized for training models < 3B
  • Supports vision models (SAM, CLIP)
Fileserver · ≥ 150 TB

Storage & Master Node

A master node and fileserver equipped with ≥ 150 TB of high-speed persistent storage to manage large datasets, model checkpoints, and evaluation runs.

  • ≥ 150 TB persistent storage capacity
  • Centralized checkpoint registry
  • Secure, high-availability backups
Up to 200 Gbit/s

InfiniBand Network Interconnect

High-speed InfiniBand network interface running up to 200 Gbit/s to minimize latency and maximize transfer rates during distributed training operations.

  • 200 Gbit/s maximum throughput
  • Low latency multi-GPU sync
  • Fast dataset streaming from fileserver

Cluster Queues

Partition layouts and nodes

Slurm uses partitions to manage resource allocations and group nodes.

All Cluster Nodes

batch (Default)

Non-interactive partition

  • Default queue for sbatch jobs
  • Supports long-running training tasks
  • Max duration up to 36 hours (llm-research)
RTX 6000 Pro Nodes Only

interactive

Interactive partition

  • Enables srun --pty bash or salloc sessions
  • Strict 6-hour time limit per session
  • Ideal for debugging and GPU checking
Gateway Node (Rechenzentrum)

Login / Frontend

cloud-201.rz.tu-clausthal.de

  • Located at the Computing Center of TU Clausthal
  • For editing code, data upload, and scheduling
  • Strictly no compute workloads allowed
01

Submit

Submit jobs via the Slurm scheduler using sbatch (batch) or srun (interactive).

02

Monitor

Monitor active jobs and resource allocations with squeue and sinfo.

03

Execute

Jobs run in isolated environments with resources allocated automatically.

Policies

Access, queue, and usage policies

Use of the Kaiser I GPU Cluster is governed by account groups and limits designed to share compute resources fairly.

Workshops & Open Lab

  • Regular workshops presenting the infrastructure to science, start-ups, and industry.
  • Active initiation of collaborative research and engineering projects.
  • Workshops registration opens after cluster commissioning (by Sep 1, 2025).

Innovation Hub

  • Connects regional universities and enterprises to strengthen technology transfer.
  • Directly supports the Lower Saxony RIS3 innovation strategy.
  • Enables resource-efficient design, circular production, and cognitive robotics.

User-Friendly Access

  • Training courses, direct hotline, and on-site support in German.
  • Full maintenance and replacement part service guaranteed for five years.
  • For collaboration inquiries, contact bartelt@isse.tu-clausthal.de.

Need detailed usage instructions?

Learn how to connect via SSH, run interactive pseudo-terminals, submit batch jobs, and manage your conda/singularity environments in our comprehensive cluster documentation.

Access Cluster Documentation
Publications

Research outputs from the network.

Papers, preprints, and workshop contributions produced using CORE Network infrastructure.

View full publication archive