Please Wait...
Job Summary:
We are looking for a backend engineer to build the Go services behind the k0rdent-ai platform — the multi-tenant control plane for enterprise GPU infrastructure. As AI workloads scale globally, you will collaborate with engineering teams across the US, Europe, and India to design and develop the API-first services that power this ecosystem.
You will own features end to end — API contract, service, schema, orchestration, tests, and packaging — with senior engineers alongside you for design review and support. We do not expect you to arrive knowing our whole stack; we expect strong Go and API fundamentals, and the appetite to learn the rest quickly. Working within an agile framework, you will directly impact the reliability and correctness of the infrastructure our customers run in production.
Main Responsibilities:
Build Go microservices, model database schemas, and write migrations for the multi-tenant k0rdent AI platform.
Design public REST APIs with OpenAPI, along with multi-language SDKs.
Implement durable orchestration in Temporal — idempotency, retries, timeouts, determinism constraints.
Build event-driven, asynchronous communication between services on a distributed streaming or message-broker platform.
Implement authentication and authorization boundaries across an enterprise IAM platform and an API gateway — OIDC/JWT, service-to-service auth, tenant-scoped RBAC, gateway routing, and rate limiting.
Contribute to the identity model — standards-based user and group lifecycle provisioning (SCIM), and unifying authorization between the platform's API permissions and Kubernetes-native RBAC.
Write unit, integration, and end-to-end tests as part of every change.
Instrument services for production — structured logging, metrics, and audit trails.
Package and ship services with Docker, Helm, and Kubernetes.
Review peers' code and debug incidents to root cause with regression coverage.
Required Skills/Abilities:
Bachelor's degree in Computer Science, Software Engineering, or a related field — or equivalent demonstrated ability.
Solid CS fundamentals: data structures, algorithms, concurrency, networking, and how HTTP and databases actually work.
4-5 years of professional software engineering, including 2-3+ years building production backend services in Go (4+ years of Go preferred).
Strong Go fundamentals: concurrency, context propagation, error wrapping, stdlib HTTP routing (Go 1.22+).
Experience building and shipping REST APIs, and a working understanding of versioning and backward compatibility.
Solid relational database skills — schema design, migrations, transactions, and query tuning.
Testing discipline: unit and integration tests are part of your definition of done, alongside code review and CI/CD (GitHub Actions or equivalent).
Clear written English for design docs, review comments, and asynchronous collaboration across global time zones.
Authorization to work in the United States.
Tech Stack (Prior Experience Preferred, Not Mandatory)
Candidates with strong Go and API design basics are evaluated primarily; prior familiarity with these tools is beneficial, though on-the-job mastery is anticipated.
Durable Workflow Execution: Temporal orchestrations.
Cloud-Native & Kubernetes Ecosystem: Container orchestration using Kubernetes, Cluster API, custom controllers and operators, Docker runtime, and deployment packaging through Helm charts.
Identity Federation & Ingress Gateway: Keycloak architecture (realms, client configurations, federated providers, auth token lifecycles) alongside Envoy or comparable edge proxies (traffic management, filter chains, rate limiting policies).
Persistence & Event Infrastructure: Relational data persistence in PostgreSQL with database schema migration workflows, coupled with distributed streaming platforms like Apache Kafka or NATS JetStream messaging.
Must Have
The "I'll figure it out" instinct — you research, experiment, and read source code before asking to be unblocked, and you ask a well-formed question when you are.
Something you built and can walk us through in detail: what it does, why you designed it that way, and what you would change now.
Willingness to be reviewed rigorously and to treat feedback as the fastest way to get better.
Testing instinct — you can explain how you know your own code works.
Authorization to work in the United States.
Nice to Have
Multi-tenant SaaS or PaaS platform experience.
Exposure to GPU infrastructure or AI/ML workload scheduling on Kubernetes.
SCIM provisioning, or federating platform permissions with Kubernetes RBAC.
sqlc, OpenAPI codegen, or other schema-driven toolchains.
Observability: Prometheus, Grafana, OpenTelemetry, audit logging.
Python or TypeScript for E2E suites and client SDKs.
Security engineering: threat modeling, secret management, dependency and container scanning.
What does offer you?
We are a Leader for Container Management in G2 (#2 after AWS)!
is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, ensures that customers retain full control of their infrastructure strategy.
Explore more opportunities that might be a good fit for you.
Get real-time job updates, apply on the go, and manage your profile easily with our mobile app.
Get it on Google Play