UC Berkeley
I'm a senior studying data science at UC Berkeley, though most of my interests are in systems and AI/ML infrastructure.
My first internship was on SageMaker HyperPod's inference team, where I built intelligent routing for the load balancer in our inference operator. I started with sticky-session routing, pinning each conversation to a single pod so follow-up requests could reuse the KV cache already sitting there, and used it to measure how much prefill work cache reuse actually saves. That work fed into prefix-aware routing, where the router tracks which prefixes are cached on which endpoint and sends each request to the pod with the best match, weighed against its current load.
I also just finished an internship on the HyperPod Dataplane team, where I built a full CI/CD pipeline for a binary artifact distribution system and shipped it to production. For our team specifically it lets us ship our agent binaries straight to a customer cluster node, without requiring them to stop training jobs.
Apart from work, I also love learning new things on the side. Right now I'm working through Build a Large Language Model From Scratch, and learning Rust. I also like DJing. Some of my favorite artists right now are KETTAMA, John Summit, and i_o.
Sponsorship Director at Cal Hacks. Reach out at collin@hackberkeley.com.
USAF Veteran, Aerospace Propulsion
Testing whether injecting live X data ahead of requests can optimize inference through KV cache reuse. Built a Rust gateway that proxies to xAI and injects shared context prefixes, with a Python benchmark suite to measure it.


Like a producer reusing loops they have made before, mainstacks lets you extract patterns from past projects and drop them into new ones. Your agents get context on how you build things, so they stop guessing and start building the way you would



Wrap your LLM client with one line of Python and every call, token, cost, and source location shows up in your dashboard in real time. Built to make agentic AI costs visible per agent, per function, and per line of code before they surprise you.



Co-authored a paper accepted to ICML 2026 AI for Science Workshop: developed a multi-agent system (MAS) that orchestrates expert decision trees with Vision-Language Models (VLMs) for automated bias labeling in forestry remote sensing, outperforming supervised ML baselines while preserving interpretability



Automates medical referral intake, OCR to structured extraction to EMR auto-fill. Taught me more about customer discovery and what people actually pay for in the context of building a startup.




My Neovim configuration. Clean, fast setup focused on development with LSP support, fuzzy finding, and Treesitter. Uses Packer.


One of my first projects, built with some friends at the first AI Hackathon from UC Berkeley. Analyzes news article bias with GPT-4 scoring and Hume AI sentiment, then rewrites the piece from three perspectives.


Working through Sebastian Raschka's Build a Large Language Model From Scratch, then building tinykv on top of it. A minimal transformer inference engine in C++ focused on the KV cache manager, the scarce resource that determines throughput in distributed LLM serving.

4.6, amazing worldbuilding, and ideas of politics, religion, control, etc
3.9, great, but definitely not as good as the first
4.0, Again, but good
4.9, probably my favorite book
3.2, fun story, nothing too crazy
3.5, remember it being decent
4.1/5 Loved this book, there were some moments where the science dragged on a bit, but overall amazing
1st book 4.3/5, 2nd 3.9, 3rd 3.3, great series though
My childhood