What I Build
10+ years building distributed, high-scale systems across adtech, fintech, and e-commerce — including roles at Google, Ant Group, and Charles Schwab.
Distributed Systems
Event-driven architectures, high-scale services, reliability and failure recovery.
Financial Infrastructure
Systems where correctness, availability and operational safety matter.
AI Engineering
RAG, evaluation, agents and AI-powered developer workflows.
Technical Leadership
Architecture, cross-team technical direction and engineering strategy.
My Journey
Tech Stack
Java, Python, GCP, AWS, Gen AI, RAG, Kafka, RabbitMQ
Professional Profile
I build distributed systems where reliability, scale and operational correctness matter.
My career spans Google, Ant Group, and Charles Schwab, working across advertising, payments, and financial infrastructure.
More recently, I've been exploring how RAG, agents, and evaluation systems can change software engineering itself.
What I do:
- Design and build scalable, production-grade distributed systems
- Drive reliability, performance, and operational safety
- Lead architecture and cross-team technical direction
Architecture Deep Dives
Designing an Ad Click Aggregation Platform at 10B+ Events/Day
Read full case study →Designing a Coding Judge System for Secure, Sandboxed Code Execution at Scale
Read full case study →Designing a Ticket Booking System for Flash-Sale Concurrency Without Overselling
Read full case study →Case Studies
A closer look at two representative problems: keeping financial infrastructure reliable, and building AI as an engineering system rather than a chatbot wrapper.
Reliable Financial Infrastructure
Problem
Financial systems require high availability, controlled deployments, event-driven communication, and strong operational guarantees.
What I worked on
- Event-driven architecture
- RabbitMQ / messaging
- Distributed services
- Failure recovery
- Production incident remediation
- Deployment and DR
- Governance and operational reliability
Impact
Designed and drove the event-driven paths, recovery behavior, and incident-response process across a set of tightly coupled distributed services — the kind of operational discipline financial infrastructure demands by regulation as much as by good practice. That meant building in recovery for when something breaks rather than assuming it won't, and being one of the people paged when production incidents needed direct remediation.
AI-Powered Engineering Systems
Problem
Engineering change review (CRQs, Jira issues, architecture references, standards) is manual, slow, and inconsistent — the evidence a reviewer needs is scattered across several systems.
What I built
Engineering Change Review Agent — an agentic workflow that retrieves Change Requests, Jira issues, architecture references, and engineering standards, then synthesizes them into an evidence-backed review.
Impact
Turns a manual, cross-system lookup process into a single orchestrated workflow that hands a reviewer an evidence-backed summary instead of a blank search bar — treating AI as an engineering system with retrieval, synthesis, and review stages, not a single API call to a chat model.
Projects
The project I'm actively building right now.

RAG Chatbot
Customer support chatbot over policy PDFs using retrieval-augmented generation, with an evaluation harness that benchmarks and tunes retrieval (Python, Streamlit, FAISS, Ollama + Gemma3). Builds on an earlier prototype, Finwise.
Let's connect!
Open to Staff and Senior engineering conversations. Happy to talk distributed systems or AI engineering.
- Dallas, TX
- Let's chat
- [email protected]
- www.bhpham.com



