Senior Software Engineer, AI Job Orchestration
Firmus Technologies
| Company | Firmus Technologies |
| Category | Engineering |
| Location | Sydney |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 24 Apr 2026 |
| Last verified | 11 Aug 2026 |
| Source | Employer ATS (greenhouse) |
Description
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Role Summary
As a Senior Software Engineer on the AI and Applications team, you'll own the control plane that powers AI workload submission across Firmus AI Platforms. You'll design and build unified job submission APIs, CLI, and web interfaces for training, inference, and fine-tuning workloads on Kubernetes and Slurm—implementing RBAC, multi-tenant isolation, resource quotas, and intelligent scheduling policies (priority classes, pre-emption, fairness). You'll create template catalog for pre-built training and inference recipes, wire observability pipelines for per-job GPU metrics cost tracking and expose telemetry APIs for platform monitoring. This role requires deep Kubernetes and Slurm expertise, strong distributed systems knowledge, and close collaboration with infra, platform, and LLM engineering teams to deliver a seamless, production-grade job orchestration experience for hyperscaler customers.
Key Responsibilities
Design and build unified job submission APIs, CLI, and web UI for all AI workload types (training, inference, fine-tuning) on Kubernetes and Slurm with Firmus AI Factory context (tenant isolation, resource requests, metadata tagging, observability hooks).
Implement comprehensive job metadata models and schemas: track job ID, job type, tenant, user, resource requirements, priority class, timestamps, lineage, execution status.
Integrate authentication/authorization (RBAC) and resource quotas; enforce multi-tenant isolation at submission time across all job types.
Build AI job scheduling and orchestration layer: priority classes, preemption policies, fairness algorithms, resource quota enforcement, and intelligent job routing.
Build the AI Factory template catalog: discovery, parameter validation, and manifest generation for training templates, inference serving templates, and fine-tuning recipes.
Wire job submissions to observability pipeline: inject labels/annotations (job_id, tenant, user, model_name, job_type) so metrics are tagged per-job.
Expose job-level telemetry APIs (GPU metrics, cost accrual, MFU progression for training; latency, throughput, tokenomics for inferencing) for platform telemetry and monitoring.
Extend job submission to handle inference workloads: design inference job specifications (model, batch size, latency SLA, cost constraints); integrate with inference serving APIs.
Coordinate with platform team on observability dashboard integration, with LLM engineers on template design, and with ModelOps on reliability standards.
Skills & Experience
5–7 years of