ResumeSoftware EngineerSeattle, WA
Nitesh Shantha Kumar
Software engineer who designs event-driven backend systems: Kafka pipelines, PostgreSQL that stays fast under load, the reliability work that keeps both running, and AI agents for engineering teams.
- Now
- Software EngineerNordstrom · Pricing & Promotions · since 2021
- Focus
- Distributed systemsEvent-driven design · data · reliability
- Toolbox
- Java · PythonKafka · PostgreSQL · AWS · Kubernetes
Read this first
- Event-driven systems on Kafka: outbox publishing, compacted topics, and schemas that evolve safely.
- Database performance: query plans, indexing, and concurrency-safe updates in PostgreSQL.
- Safe change: feature flags, shadow mode, parity checks and rollback plans before any cutover.
- Technical lead who also builds agentic tooling for incident triage.
01
What I know
The concepts and tools I work with, grouped by area. Open one to see what’s inside.
Event-driven systemsDesigning Kafka pipelines that stay correct when things fail.Kafka
- Transactional outbox
- Write state and the event in one database transaction; a relay publishes and marks rows sent only after the broker confirms. See Fig A.
- Delivery semantics
- At-least-once delivery with idempotent consumers, and keys chosen to keep per-entity ordering.
- Log compaction
- Keep the latest value per key, so a topic can seed and refresh a local read store.
- Schema evolution
- Avro schemas with compatibility rules, so producers can change without breaking consumers.
- Ownership & access
- Topic ACLs, additive permission changes, and Confluent Cloud.
Databases & performanceReading query plans and fixing the cause, not the symptom.PostgreSQL
- Query plans
EXPLAIN ANALYZE, sequential vs index scans, and what the planner’s cost estimates mean.- Sargable predicates
- Wrapping an indexed column in a function such as
COALESCEquietly turns off the index. - Concurrency control
- Atomic conditional updates instead of read-check-write races; row locks and isolation levels.
- Maintenance
- MVCC, dead tuples, autovacuum, table bloat, and transaction-ID wraparound.
- Schema changes
- Versioned Flyway migrations with an undo path.
- Beyond Postgres
- Oracle, DynamoDB, Neo4j graph modelling, and Redis / Valkey caching.
System designChoosing trade-offs on purpose, and writing them down.Architecture
- Sync vs async
- When a call belongs on the request path, and when the work should be queued.
- Backpressure & rate limits
- Protect downstream services by sizing throughput from measured rates. See Fig B.
- Batching & timeouts
- Bounded batches, per-call timeouts, and commit-and-continue, so one failure doesn’t stall the rest.
- Caching with fallback
- Event-fed local read stores, with the source service as the fallback on a miss.
- Decision records
- ADRs that name the rejected alternative and the constraint that ruled it out.
Reliability & operationsShipping changes safely, and knowing when something breaks.On-call
- Safe migrations
- Feature flags split by stage, shadow mode in production, parity checks, and tiered rollback plans.
- Alerts that can fire
- Alerting on missing signals and heartbeats as well as errors, with triage steps for every alert.
- Incident response
- Contain first, then fix the structure; blameless RCAs with causal chains and follow-ups.
- Load testing
- JMeter tests past the expected peak before launch.
- Observability
- New Relic, Splunk, and custom metrics for throughput, latency and failures.
Cloud & platformRunning services on AWS and Kubernetes.AWS
- AWS
- EKS, S3, RDS and Aurora, including right-sizing clusters once load tests are done.
- Kubernetes & Docker
- Deployments, cron jobs, and how autoscaling multiplies load on shared resources like connection pools.
- Delivery
- CI/CD, environment-driven feature flags, and deploy and rollback runbooks.
AI & developer toolingBuilding agents that do real engineering work.Agents
- Agentic workflows
- Skills and subagents that gather context across chat, observability, code and ticketing tools.
- Model Context Protocol
- MCP servers that give models controlled access to databases and tools.
- Orchestration
- LangGraph ReAct agents, and natural language to SQL or Cypher with automatic retries.
- Retrieval
- Indexing code into knowledge graphs with vector embeddings.
FoundationsRobotics research and safety-critical software.Robotics
- Robotics
- ROS, motion planning, and perception and sensor fusion for autonomous vehicles.
- Safety-critical software
- DO-178B/C verification and validation, with test harnesses run on target hardware.
- Languages
- Java, Python, SQL, C++ and Ada.
Fig A · The transactional outboxAll steps
Fig B · Throughput ceilingBack-of-envelope
Ceiling per cycle350 rows
Share of queue per cycle1.8%
Time to drain at ceiling48 min
02
Experience
Nordstrom
Seattle, WA Software Engineer — Pricing and Promotions PlatformAug 2021 – Present- Technical lead for a clearance and promotions platform with REST and Kafka ingestion. I split delivery into parallel workstreams and own design review and integration.
- Move legacy batch pipelines to event-driven publishing on Kafka with the transactional outbox pattern, using feature-flagged, shadow-mode cutovers.
- Design for scale and resilience: asynchronous processing, rate limiting, event-fed read stores, and fallbacks to upstream services.
- Diagnose database and throughput problems, lead incident investigations, and write root-cause analyses.
- Build alerting and runbooks, plus agentic tooling that automates incident triage.
Ak Aerotek
Bengaluru, India Software EngineerMay 2016 – May 2019- Verification and validation of safety-critical avionics software to DO-178B/C, writing Ada test harnesses and running them on target hardware.
- Led a small team through code and audit reviews.
03
How I work
Habits I bring to every system I touch.
- Baseline before designing.Measure latency, error rate and volume in production before proposing a change.
- Write down the road not taken.Every decision record names the rejected alternative and the constraint behind it.
- Prove parity before cutover.Shadow mode, split feature flags, and a rollback plan for each named failure.
- Contain first, then fix the structure.Stop the bleeding quickly, then remove the cause so it can’t recur.
- Prove a fix is safe, not just fast.Show the output is identical before shipping an optimization.
- Alerts ship with triage steps.An alert nobody knows how to act on is only half an alert.
04
Skills
Chips marked ↳ open the matching area in What I know.
Languages
JavaPythonSQLC++
Distributed systems
AvroLog compactionConfluent Cloud
RESTful APIs
Data
OracleDynamoDBNeo4jRedis / Valkey (ElastiCache)
Cloud & operations
KubernetesDockerNew RelicSplunkJMeterFlyway
AI & tooling
MCPLangGraph
05
Education
M.S., Engineering Sciences (Robotics) — University at Buffalo, NY
2019 – 2021Graduate Research Assistant in the LANDR Lab, working on autonomous-vehicle perception and sensor fusion.
B.E., Electronics & Communication — Visvesvaraya Technological University, India
2011 – 2015Let’s talk systems, agents, and the engineering between them.
Email is the best way to reach me.
nithesh1001@gmail.com