Data platforms
Modernizing Spark infrastructure, helping build Apache Iceberg infrastructure, and making a $6–7M/year Spark fleet’s costs visible.
Yiping Deng
Tech Lead II · AI Data · HubSpot
I build the data platforms behind AI at HubSpot — from distributed systems and vector search to GPU infrastructure and observability.
01 / What I build
Reliable platforms. Thoughtful abstractions.
Room for the next big idea.
Modernizing Spark infrastructure, helping build Apache Iceberg infrastructure, and making a $6–7M/year Spark fleet’s costs visible.
Scaling vector search to 7B+ production ANN vectors across 200+ indices, and building inference and feedback pipelines that ingest 2 TB/day.
Building OpenTelemetry tracing and storage for all AI agents at HubSpot, and using AI agents to analyze data-entity lineage graphs.
Selected projects / GitHub
A few things I’ve built outside the day job — from GPU experiments and developer tools to useful little systems for everyday life.
A hands-on GPU programming notebook: CUDA matrix-multiplication kernels, Triton exercises, and custom PyTorch extensions.
Mock shell-tool responses and verify expected commands with reusable fixtures, exact or regex matching, and pytest helpers.
A Rust CLI for exploring Parquet files on S3, including folders, SQL queries through DataFusion, and AWS profile support.
Turn a Git repository into structured AI context, with a file tree, source contents, include/exclude patterns, and .gitignore support.
An e-ink dashboard for LLM usage and local coding-agent activity, with touch navigation, cached views, and explicit stale-data states.
A Python CLI for CatLink smart litter boxes: device status, logs, cleaning controls, and authentication backed by the system keyring.
02 / The journey
HubSpot / Dublin, Ireland
HubSpot / Dublin, Ireland
HubSpot / Dublin, Ireland
HubSpot / Dublin, Ireland
HubSpot / Dublin, Ireland
Amazon / Luxembourg
Jacobs University Bremen / Bremen, Germany
Ubimax / Bremen, Germany
03 / A little about me
My work sits where data infrastructure meets machine learning: the systems that collect, store, retrieve, and make sense of data at scale. At HubSpot, I’ve grown from an AI infrastructure engineer into a tech lead, building platforms and teams along the way.
I’m also a contributor to open-source infrastructure and a co-author of research in formal theorem proving. From mathematical foundations to production systems, I like understanding how things work — and making them work better.
Find me on GitHubMinor in Intelligent Systems
Jacobs University Bremen
Coursework in machine learning, algorithms, data structures, operating systems, and computer vision. Research in formal logic and marine robotics.
04 / Beyond the day job
Four merged contributions to the open-source vector database: peer mTLS, HTTPS client-certificate validation, live TLS certificate rotation, and clustered-collection regression tests.
Explore the contributionsCo-author · EasyChair Preprint 152
Read the publication05 / The notebook
Notes on code, computer science, and the ideas underneath. From the archive.
Mixing data-fetching behavior into React components with a reusable higher-order component.
A hands-on introduction to map, reduce, and counting characters with Haskell.
A look at string searching, the bad-character rule, and the good-suffix rule.
Good things start with a conversation.
Talk systems, trade ideas, or just say hello.