Hey, I'm Bill Pham.

10+ years building systems at scale
Charles Schwab — Financial Infrastructure Google — Ads Infrastructure Ant Group — Alipay / Financial Systems
What I Do

What I Build

10+ years building distributed, high-scale systems across adtech, fintech, and e-commerce — including roles at Google, Ant Group, and Charles Schwab.

Distributed Systems

Event-driven architectures, high-scale services, reliability and failure recovery.

Financial Infrastructure

Systems where correctness, availability and operational safety matter.

AI Engineering

RAG, evaluation, agents and AI-powered developer workflows.

Technical Leadership

Architecture, cross-team technical direction and engineering strategy.

About

My Journey

Bill Pham

Tech Stack

Java, Python, GCP, AWS, Gen AI, RAG, Kafka, RabbitMQ

Certificates

GCP Cloud Architect badge GCP Data Engineer badge AWS badge GCC Data Analytics badge

Social

Professional Profile

I build distributed systems where reliability, scale and operational correctness matter.

My career spans Google, Ant Group, and Charles Schwab, working across advertising, payments, and financial infrastructure.

More recently, I've been exploring how RAG, agents, and evaluation systems can change software engineering itself.

What I do:

  • Design and build scalable, production-grade distributed systems
  • Drive reliability, performance, and operational safety
  • Lead architecture and cross-team technical direction
Let's connect: [email protected]
Deep Dives

Architecture Deep Dives

Designing an Ad Click Aggregation Platform at 10B+ Events/Day

Read full case study →

Designing a Coding Judge System for Secure, Sandboxed Code Execution at Scale

Read full case study →

Designing a Ticket Booking System for Flash-Sale Concurrency Without Overselling

Read full case study →
Engineering

Case Studies

A closer look at two representative problems: keeping financial infrastructure reliable, and building AI as an engineering system rather than a chatbot wrapper.

Reliable Financial Infrastructure

Problem

Financial systems require high availability, controlled deployments, event-driven communication, and strong operational guarantees.

What I worked on
  • Event-driven architecture
  • RabbitMQ / messaging
  • Distributed services
  • Failure recovery
  • Production incident remediation
  • Deployment and DR
  • Governance and operational reliability
Impact

Designed and drove the event-driven paths, recovery behavior, and incident-response process across a set of tightly coupled distributed services — the kind of operational discipline financial infrastructure demands by regulation as much as by good practice. That meant building in recovery for when something breaks rather than assuming it won't, and being one of the people paged when production incidents needed direct remediation.

AI-Powered Engineering Systems

Problem

Engineering change review (CRQs, Jira issues, architecture references, standards) is manual, slow, and inconsistent — the evidence a reviewer needs is scattered across several systems.

What I built

Engineering Change Review Agent — an agentic workflow that retrieves Change Requests, Jira issues, architecture references, and engineering standards, then synthesizes them into an evidence-backed review.

Engineer
Agent
CRQ retrieval
Jira retrieval
Confluence retrieval
Standards / policies
Evidence synthesis
Risk / compliance review
Impact

Turns a manual, cross-system lookup process into a single orchestrated workflow that hands a reviewer an evidence-backed summary instead of a blank search bar — treating AI as an engineering system with retrieval, synthesis, and review stages, not a single API call to a chat model.

Work

Projects

The project I'm actively building right now.

RAG Chatbot preview

RAG Chatbot

Customer support chatbot over policy PDFs using retrieval-augmented generation, with an evaluation harness that benchmarks and tunes retrieval (Python, Streamlit, FAISS, Ollama + Gemma3). Builds on an earlier prototype, Finwise.

Writing

From the Blog

Technical write-ups land here first; older pieces are archived on Medium.

Building an Evaluated RAG System: From Naive Retrieval to Measured Quality thumbnail

Building an Evaluated RAG System: From Naive Retrieval to Measured Quality

A naive RAG pipeline retrieves something and an LLM answers. The real engineering work is proving a change to that pipeline actually helped — with numbers, not a vibe check.

Read more →
Get In Touch

Let's connect!

Open to Staff and Senior engineering conversations. Happy to talk distributed systems or AI engineering.