---
title: "Senior Backend Engineer, Vision at Sarvam"
canonical: https://careercto.dev/jobs/senior-backend-engineer-vision-sarvam
---

# Senior Backend Engineer, Vision at Sarvam

> Bengaluru, India · on-site · full-time

## About the role

You will own the architecture of the serving harness for Sarvam's vision models — the system that has to deliver frontier-grade extraction quality out of 3B and 30B in-house models, at national scale, with cost and latency budgets that actually close\. This means owning the hard trade-off surface directly: accuracy versus latency versus rupees per page\.

## What you'll do

Own the end-to-end architecture of the OCR and extraction serving harness: API layer, orchestration, inference layer, post-processing, delivery Design the accuracy harness — multi-pass extraction, ensembling, cross verification, schema-constrained decoding, confidence calibration, targeted re runs — and prove its gains against held-out evaluation sets Architect durable, resumable document workflows in Temporal: fan-out across pages, partial failure recovery, exactly-once side effects, long-running jobs measured in minutes to hours Own the inference serving layer alongside infra: batching strategy, GPU pool management, autoscaling on real signals, queue depth and admission control, multi-model routing

## Looking for

5–6\+ years in backend engineering, with meaningful time spent operating high throughput production systems you were on-call for Deep proficiency in Go and/or Python, and the judgement to know which belongs where Strong distributed systems design: queues, workflow orchestration, idempotency, backpressure, retry and timeout semantics, consistency trade-offs, graceful degradation Production experience with Temporal or an equivalent durable execution engine, on workflows that mattered Kubernetes in production — autoscaling, resource management, rollouts, debugging under load; GPU workload scheduling is a strong plus Demonstrable experience serving ML or LLM inference in production: batching, caching, model versioning, A/B rollout, latency budgeting

## Nice to have

Direct experience with OCR, IDP, or document AI systems — Textract, Document AI, Azure DI, or something you built yourself GPU inference stacks: vLLM, TensorRT-LLM, Triton, SGLang, Ray Serve Evaluation infrastructure for ML systems — golden sets, regression gates, human in-the-loop review loops Experience with BFSI, healthcare, or public-sector compliance and data-residency constraints On-prem or air-gapped deployment experience Note We are looking for people who can own the outcomes described here, not people who match every line of this specification\.


## Details

- Skills: Python, Azure
- Published: 2026-08-18T00:00:00Z
- Apply: https://jobs.ashbyhq.com/sarvam/3652221b-d408-486b-9ad1-ab073aaf1071/application
