yash backend & data engineer
WorkAboutProductsContact

Work

Projects & systems

Things I have designed or built — from a bank-scale batch data platform to local inference and privacy-first products. Filter by discipline — links work without JavaScript.

Showing 2 projects tagged “Microservices”

Batch Operations Data Platform

2022–2026

At TCS, I own the data platform behind a Fortune-500 bank's batch operations. It pulls job status out of the bank's mainframe, stores it, and powers the dashboards and alerts the client's operations team lives on. I designed the query engine, the failure tracking, and the system that predicts an SLA miss before it happens. The result: our most important batch job went from ~15 minutes to under 1 minute, and the platform now flags problems before the client ever sees them. Under the hood: Python, Oracle, and MongoDB, with a config-driven SQL engine and 85% test coverage across the codebase.

  • Cut the core batch job from ~15 minutes to under 1 (93% faster), with results verified unchanged
  • Predicts SLA misses before they occur, using job-calendar data
  • Catches silent data failures early, instead of after the client does
  • 85% test coverage, clean static analysis, and structured logging (70% less log volume)

Local LLM Inference Platform

2026

I run my own large language model on a single desktop machine — no cloud, no external API. I set up the model, the serving, and the tooling so it can be driven like any other API. On top of that I built an automated pipeline that turns 80 work records into a deduplicated, evidence-traced memory where every output carries its source. The hard part isn't the model — it's making a model that can be wrong produce output you can trust and trace. Under the hood: an NVFP4-quantized Qwen 27B model, a 262K-token context window, and an OpenAI-compatible API via SGLang (migrated from vLLM behind the same contract).

  • A 27B model running locally on one 128 GB desktop box — no cloud
  • 262K-token context, exposed as an OpenAI-compatible API
  • Automated pipeline: 80 records, zero missed after fallback, every output traceable to its source
  • Hybrid deduplication that only escalates ambiguous cases to the model

Want the unpolished version? The about page has more context on how these fit together.