Hyderabad, India · Enterprise AI platforms

The operating layer for production AI.

I'm Soumendra Kumar Sahoo. Over fourteen years I've moved from mainframes to Big Data, multi-cloud and AIOps at IBM, enterprise search, MLOps and LLMOps. Today I own the layer that decides whether an AI system survives real users: architecture, instrumentation, evaluation, governance, and the people who run it. I still read the traces myself.

Soumendra Kumar Sahoo
Currently
AI Observability Architect PepsiCo · since 2024

Fourteen years · six organisations

  • PepsiCoSince 2024
  • Freshworks2021 – 24
  • IBM2018 – 21
  • Wipro2016 – 18
  • Accenture2014 – 16
  • Tata Consultancy Services2012 – 14

The operating layer

Five stages decide whether an AI system survives production.

Most AI failures are operating failures. What breaks in production is evaluation nobody owns, instrumentation added the week after the incident, and governance written by people who have never read a trace. This is the span I work across.

01

Architect the platform

Start from the decision the organisation needs to make, the risk it needs to control and the feedback it needs to learn. Architecture follows those constraints. The goal is a shared capability several teams can use, instead of another heroic one-off every quarter.

Evidence

03

Evaluate continuously

Evaluation is a standing mechanism with an owner and a cadence, running long after launch. It also needs evaluating itself. LLM-as-judge carries real bias and a real bill, and retrieval architecture changes what "correct" even means.

Evidence

04

Govern the risk

Policy, accountability, monitoring and continuous improvement assembled into one operating system rather than a document nobody opens. Autonomy earns bounded permissions, verified backups and tested recovery, the same controls any other production system gets.

The record

Numbers you can check.

Each one is verifiable against my profile, my publications, or the archive on this site.

14+

years in enterprise engineering, 2012 to today

6

organisations, from TCS to PepsiCo

115

posts published since May 2012

OTCA

OpenTelemetry Certified Associate

Selected writing

Notes for people operating real systems.

I write to make architecture choices, leadership patterns and failure modes easier to inspect.

All writing

What failed

Every weekly note has this heading. So does this page.

Most leadership sites publish only outcomes. A standing failure column is harder to fake, and it shows how someone reasons when the system doesn't cooperate.

My research agents fabricated their sources.

Eleven invented model names and four fake paper titles, produced with total confidence. The fix was a designed verification control. Better prompting would never have caught it.

Week 20, 2026
A silent backup failure cost me my skills directory.

No alert, no verification step, nothing to restore from. Agents need bounded permissions and tested recovery like any production system.

Week 31, 2026
I failed the OTCA exam by two percent.

First attempt, strict remote proctoring. The second attempt worked because I changed the preparation method rather than adding effort.

January 2026
I built a speed reader and stopped using it in a day.

A few hours with AI tools produced a working prototype for a problem I didn't actually have. Building isn't the same as solving.

January 2026

Read the weekly notes

Let's compare notes

I'm always interested in the operating problems behind production AI.

Leadership roles, advisory, speaking, community work, or a straight argument about evaluation design. Send a note and I'll reply by email.

New here? How to work with me covers what to put in a first message.

Prefer your own client? contact@soumendrak.com