Home / Case studies

Case studies in distributed and real-time systems

Four projects from Atlassian, Egnyte and ServiceNow, told through the problem and the decisions. Internal details stay internal. Each one ends with what I'd tell another team facing the same problem.

−80%

Company
Atlassian · Commerce (checkout, billing and provisioning)
Role
Owned the end-to-end test migration and the cross-team fix effort
Stack
TypeScript, React, Node.js, GraphQL, Playwright

Cutting flaky end-to-end tests by 80%

The problem

The commerce end-to-end suite failed often enough that engineers re-ran it by reflex. A red build no longer meant a broken product, so real regressions slipped through and deploys slowed down.

What I did

  • Audited eight months of recorded failures and classified each one by root cause instead of by test name.
  • Drove the migration of the suite from Cypress to Playwright, using the move to fix waiting and isolation problems at the source.
  • Coordinated fixes with the teams that owned the underlying causes, since many failures came from shared services, not the tests.

Result

The flaky-test rate dropped by 80% and the team trusted CI again, which raised deployment confidence for every commerce release.

What I'd tell another team

Treat flakiness as a concurrency problem. Most "random" failures fall into a handful of buckets: shared test data, timing assumptions, environment dependencies, and real race conditions in the product. Count failures per bucket and fix the biggest bucket first. Quarantine a flaky test only with a named owner and a deadline, or quarantine becomes deletion.

2-way

Company
Egnyte · integration with Google Workspace
Role
Architect of the synchronization layer
Stack
TypeScript, Node.js

Real-time sync across two different data models

The problem

Users needed to collaborate on the same files from two platforms in real time. The two sides stored files differently and modelled permissions differently, so a naive copy in each direction would drift, loop, or expose files to the wrong people.

What I did

I architected a bidirectional synchronization layer between the two platforms that kept content and changes flowing both ways in real time, reconciling their storage and permission models.

Result

Real-time collaboration across both platforms, despite their different storage and permission models.

What I'd tell another team

Decide which side owns each field before writing any code. Make every apply step idempotent and version-aware so echoes of your own writes are recognised and dropped instead of looping. Permission mapping between two models is lossy, so choose the safe direction explicitly. And run a periodic reconcile job even when the event path works, because it's the only thing that catches the drift you didn't predict.

−30%

Company
Egnyte · enterprise content platform
Role
Designed and built the caching and retrieval layer
Stack
Node.js, Redis

A caching layer that cut query latency by 30%

The problem

Queries for large enterprise customers were slow, and retrieval paths were repeating the same expensive work.

What I did

Engineered a Redis and Node.js caching and retrieval layer in front of the hot paths.

Result

30% lower query latency and 20% faster retrieval at enterprise scale.

What I'd tell another team

A cache is a second copy of your data, so it's a consistency decision first and a performance decision second. Invalidate on the events that change the data rather than on timers wherever you can, measure p95 and p99 rather than averages, and protect the backing store from a stampede when a popular key expires.

LLM

Company
ServiceNow · Generative AI (prompt-to-app product)
Role
Core architect for orchestration, prompt management and output validation; mentored 3 engineers
Stack
TypeScript, Node.js, LangChain

Orchestration and validation for a prompt-to-app product

The problem

Turning a plain-text prompt into a working application takes many model calls in sequence. Each call can be slow or return something malformed, and one bad step breaks everything after it.

What I did

  • Owned the LLM orchestration, prompt management and output-validation layers.
  • Built high-throughput orchestration pipelines that used parallelism and caching to cut end-to-end generation latency.
  • Set up CI/CD practices, led code reviews, and mentored three engineers on production LLM patterns and A/B testing of model configurations.

Result

Lower end-to-end generation latency and a validation layer that stopped malformed output from reaching later steps.

What I'd tell another team

Validate structure before meaning: a schema check is cheap and catches most failures early. Run independent steps in parallel and cache on normalised inputs. Treat a model configuration (model, prompt version, parameters) as a deployable artifact with its own experiments and rollback, exactly like code.

Facing something similar?

Tell me about the system and the symptom. I'll reply within one business day.

Start a conversation