Staff Engineer, DevOps


About kAIgentic

kAIgentic is building the intelligence layer for the world’s most ambitious enterprises. Headquartered in Singapore with teams in India and Japan, we help large organizations evolve as fast as technology itself by turning the tacit know-how locked inside their people into safe, governed, AI-powered operations.

The hardest part of enterprise transformation is not strategy. It is execution. Institutional knowledge lives in people’s heads, systems are fragmented, and risk tolerance is low. kAIgentic captures how work actually happens, designs better workflows, and runs them inside an intelligence layer that is observable, auditable, and engineered for the most regulated environments on earth. The outcome is an enterprise that continuously improves.

We are backed by SMBC Group as our founding partner and customer zero, and our platform is already being proven inside one of the most complex, regulated operating environments in the world. That means real problems, real data, and real production impact from Day 1.

 

The Role

As a Staff Engineer on the DevOps/MLOps team, you will be the technical authority for infrastructure and operations across kAIgentic’s enterprise platform. You’ll define how we deploy, scale, observe, and operate agentic AI systems reliably in the most demanding regulated environments.

What You’ll Do

  • Define the long-term infrastructure vision for kAIgentic’s global enterprise deployments; architect multi-region, multi-tenant cloud infrastructure meeting bank-grade reliability and compliance requirements
  • Design the unified CI/CD platform strategy spanning application code, infrastructure, ML models, and RAG pipelines; own the Kubernetes platform architecture including GPU federation, cost optimization, and enterprise isolation
  • Define the observability architecture unifying application metrics, agentic workflow tracing, LLP performance monitoring, and audit logging; architect the LLM operations strategy including multi-provider routing, failover, cost optimization, and latency SLOs
  • Design the MLOps platform roadmap for model lifecycle management at enterprise scale; establish the reliability engineering culture including SLO frameworks, chaos engineering, and incident management
  • Define infrastructure security architecture including zero-trust networking, secrets management, and compliance automation; drive platform efficiency and cost optimization across all AI workloads; influence infrastructure direction across all engineering teams
  • Evaluate emerging technologies in cloud-native and AI infrastructure; mentor leads and senior engineers on platform architecture; represent kAIgentic’s infrastructure capabilities to enterprise customers and support deployment planning

What You’ll Bring

  • 12+ years of experience in DevOps, SRE, or platform engineering; AI-native velocity as a default mode of working (mandatory)
  • Expert-level proficiency in Python and Go; deep expertise in 4+ of the following:; cloud architecture for regulated enterprises (multi-region, compliance)
  • Kubernetes platform engineering at scale; CI/CD and GitOps architecture for complex systems; observability architecture (metrics, traces, logs at scale); reliability engineering and SRE practices; LLM operations (serving, routing, optimization)
  • Infrastructure security and compliance automation; cost optimization for AI/ML workloads at scale
  • Experience operating infrastructure for financial services or similarly regulated industries; track record of defining infrastructure strategy adopted across organizations
  • Exceptional architectural judgment and communication skills; ability to translate infrastructure capabilities for business stakeholder

 

Why join kAIgentic?

We are a global team of builders who thrive in ambiguity, care deeply about the customers we serve, and believe the intelligence layer is how enterprise work will be reshaped over the next decade. We are building the connective tissue that lets large companies operate with the speed of a startup and the trust of an institution.

We look for people who:

  • Combine technical excellence with genuine customer empathy.
  • Are entrepreneurial and energized by zero-to-one problems with no playbook.
  • Lead with ownership, integrity, and collaboration, not titles.
  • Want to help define a new category of enterprise AI, not just ship inside an existing one.

Working here means being surrounded by peers who challenge assumptions, celebrate progress, and build with both courage and care.

 

Life at kAIgentic

  • Intelligence layer at the core. You will be building the substrate that turns institutional knowledge into governed, production-grade operations. This is not a wrapper on a model. It is enterprise infrastructure with real consequences.
  • Innovation at enterprise scale. Startup velocity meets the depth, scale, and stakes of mission-critical, regulated environments. Both are non-negotiable.
  • Ownership from Day One. Your work directly shapes the product, the culture, and the outcomes our customers see.
  • Learning and growth. You will work alongside seasoned leaders from leading enterprises who have built and scaled global businesses.
  • A culture of trust. Psychological safety, transparent disagreement, and disciplined experimentation are how we operate, not slogans on a wall.
  • Global collaboration. Teams across Singapore, India, Japan, Europe, and the US, working as one.
  • A mission worth the effort. Building something the world has not seen before: an intelligence layer that helps enterprises continuously improve how they run.
Apply Now