CosmicAC Logo

CosmicAC jobs

What a job is in CosmicAC, the job types, and the statuses a job moves through.

A job is a unit of work that CosmicAC runs on your Kubernetes cluster. When you create a job, CosmicAC provisions the resources that the job needs, runs the workload on one or more GPUs in your cluster, and tracks the job's status. Each job belongs to the team that's active when you create it. See Teams and roles.

Job types

CosmicAC runs two types of jobs. Both run inside a KubeVirt virtual machine instance (VMI) with one or more GPUs, but you work with each type differently.

  • GPU Container Job: gives you shell access to the VMI, so you can use it like a remote machine and run your code. See GPU Container Job.
  • Managed Inference Job: serves an open source model behind an OpenAI-compatible API. It uses the vLLM runtime for language models or the Parakeet runtime for speech-to-text. See Managed Inference Job.

Job status

A job moves through a sequence of statuses from the moment you create it. Each status reflects the job's current state on your cluster. The web interface and the CLI show the following statuses.

  • Creating: validating the job's configuration and GPU requirements.
  • Queued: allocating the GPUs that the job needs on your cluster.
  • Pending: waiting for CosmicAC to place the job. It holds no resources yet.
  • Starting: pulling the job's image and setting up the job. For a Managed Inference Job, this includes loading the model weights.
  • Active: the job is up and accessible. A GPU Container Job provides shell access, and a Managed Inference endpoint accepts requests. The API reports this status as running.
  • Degraded: some of the job's replicas failed while at least one is still running or starting, so the endpoint stays up on the replicas that remain. A job that runs a single replica moves to Failed instead.
  • Restarting: replacing the job's VMI while keeping its storage and resources.
  • Failed: the job failed to start or crashed.
  • Unknown: the job has no state that CosmicAC can report.

A job status of Degraded means some of the job's replicas failed. A model health value of Degraded means the replicas that remain answer requests poorly. CosmicAC reports the two separately, so an Active job can report Degraded model health.

Job logs

A job runs inside a VMI on your cluster. Its logs record what happens there.

CosmicAC publishes those records to cosmicac-wrk-monitor, which streams them live and passes them to your Loki instance. The Logs tab on a job page reads the live stream and the history in Loki. Without Loki, the tab shows live logs only. See Observability architecture.

Each record carries a source, because two parts of CosmicAC write records.

  • System: how CosmicAC sets the job up, including the steps that create its VMI and whether each one succeeded. Both job types produce these. If a job never reaches Active, these records show how far it got.
  • Application: what runs inside the job once it starts. Only a Managed Inference Job produces these, because CosmicAC starts the model server itself and captures its output through the inference agent.

Job actions

After you create a job, you can take the following actions on it.

  • Restart: replaces the job's VMI but keeps its storage and resources.
  • Delete: removes the job's VMI, resources, and storage.
  • Duplicate: opens the new job form with the job's configuration filled in. Only the web interface has this action.

For a Managed Inference Job, you can also restart or delete a single replica while it's Queued, Pending, or Starting.

For the CLI commands, see CLI commands for jobs.

Next steps

On this page