<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Eric Ma's Blog</title><link href="https://ericmjl.github.io/blog/" rel="alternate"/><link href="https://ericmjl.github.io/blog.xml" rel="self"/><id>urn:uuid:a7611166-dd1f-3792-b62b-0c03a4283350</id><updated>2026-09-11T00:00:00Z</updated><author><name/></author><entry><title>Funny Money Allocation: A Decision-Making Trick</title><link href="https://ericmjl.github.io/blog/2026/9/11/funny-money-allocation/" rel="alternate"/><updated>2026-09-11T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:bc947a1f-ea73-3dc5-9a8f-aaa0c05d41d9</id><content type="html">&lt;p&gt;A while back, I was part of a team that was going in circles. We had three options in front of us, and every discussion round ended the same way: people advocating for their preferred option in words, others pushing back in words, and no convergence in sight. There was no slam dunk benefit to any one option, hence the loop. As a team member (not the team lead), I started to feel that we needed a different way to decide than continuing to talk in qualitative terms. The Bayesian in me piped up: let's express our beliefs as probability distributions, and decide from there.&lt;/p&gt;
&lt;p&gt;I'm pretty sure I'm not the inventor of this trick; someone out there has probably formalized it under a fancier name. But I did arrive at it that day while watching my team talk past each other, and it worked well enough that I've wanted to write it up ever since. I call it funny money allocation.&lt;/p&gt;
&lt;h2 id="the-mechanics"&gt;The mechanics&lt;/h2&gt;&lt;p&gt;Here is what I did. I gave every person 1,000 points of funny money, effectively \$1,000 each, and laid out the ground rules:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;There are three options on the table.&lt;/li&gt;
&lt;li&gt;Allocate your 1,000 points across them however you like: all of it on one outcome, or any split you want.&lt;/li&gt;
&lt;li&gt;You also don't have to spend all 1,000 points. Express that you're unsure by spending fewer points and keeping the funny money for yourself.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That last rule matters more than it looks. Keeping points is really an allocation to a fourth outcome: "I don't know." It's a legitimate epistemic state, and the design makes room for it instead of forcing it to hide inside a pick. In a vote, uncertainty gets collapsed into a choice. Here, it gets its own allocation.&lt;/p&gt;
&lt;p&gt;One ground rule, baked into the setup, does all the work: your total allocation can't exceed 1,000 points. Words are unbounded; declaring "I strongly prefer option two!" costs nothing. Points are capped. The moment you have to take points away from one option to fund another, you have to rank your own convictions. And once you divide each allocation by 1,000, you have exactly what I was after: a probability distribution over the outcomes, one per person.&lt;/p&gt;
&lt;h2 id="what-the-bets-revealed"&gt;What the bets revealed&lt;/h2&gt;&lt;p&gt;When the points came in, the exercise turned out to be very revealing.&lt;/p&gt;
&lt;p&gt;I bet small: 70 points out of my 1,000, keeping the other 930. I genuinely wasn't knowledgeable about the underlying biology separating the outcomes, and my allocation said so out loud. Others bet big: 900 points on one option, leaving just 100 split between the remaining two. Where I was shrugging, they were leaning hard.&lt;/p&gt;
&lt;p&gt;That 70-versus-900 spread was exactly the information we needed. Pool everyone's allocations and you get the team's collective distribution: which option had the most support, how concentrated that support was, and who held it. The people with the deepest expertise had put their (funny) money where their mouths were, and their bets helped us see very clearly where the team's preferences lay.&lt;/p&gt;
&lt;p&gt;This is analogous to the kind of thing my former colleague &lt;a href="https://www.linkedin.com/in/clayton-springer-5a48072/"&gt;Clayton Springer&lt;/a&gt; would have prescribed. His version: state your hypothesis in the form of a molecule, because that's concrete. Ours: state your beliefs in the form of a funny money distribution, because we were also trying to be very concrete. Either way, the point is to make your belief inspectable.&lt;/p&gt;
&lt;h2 id="from-bets-to-a-decision"&gt;From bets to a decision&lt;/h2&gt;&lt;p&gt;Beliefs expressed as probability distributions made real discussion possible. Once they were on the table, the argument shifted from "my adjectives vs. your adjectives" to "why did you allocate that way?", and the second question is much more productive. The exercise also surfaced something we hadn't articulated: there was more than one legitimate way to converge.&lt;/p&gt;
&lt;p&gt;Consensus was the path we took. We talked through the distributions and rounded off some of the numbers to make the final call easier to act on and to fit our actual sample numbers. Summing up the numbers like this actually allows those with strong opinions to sway and influence the decision. And that's the point: with beliefs expressed as distributions, the decision maker can get creative. Another path is to have the person who needs to make the call treat the funny money allocation as a quantitative measure of belief, marry it with the qualitative beliefs that people have (the articulated reasons for their allocations and such), do some behind-the-scenes work, and come to the final conclusion. Done that way, everybody feels as if they've had their opinion heard, even if the outcome may not be exactly what they would have prescribed.&lt;/p&gt;
&lt;h2 id="why-this-beats-voting"&gt;Why this beats voting&lt;/h2&gt;&lt;p&gt;I think this trick has real advantages over an outright vote.&lt;/p&gt;
&lt;p&gt;A vote collapses your belief into a single choice: one person, one bit. Funny money lets me express the strength of my conviction. When I want to say "I lean this way, but only slightly", I bet 200 points instead of 900. Degree of opinion survives the aggregation instead of getting flattened into a binary.&lt;/p&gt;
&lt;p&gt;Funny money also separates confidence from preference, which voting mashes together. "Low confidence, mild preference" looks like 150 points on my preferred option and the rest in my pocket. "High confidence" looks like a big, concentrated bet. Those are genuinely different epistemic states, and they deserve different representations.&lt;/p&gt;
&lt;p&gt;Abstention, or rather deferring to experts, is built into the design of funny money allocation. Choosing not to bet on one of the three outcomes, or choosing not to bet heavily on one of the three outcomes, is a visible and legitimate stance that says, "I'm willing to defer to the experts." You can participate honestly while being genuinely uncertain, and your honesty stays visible in the pooled distribution.&lt;/p&gt;
&lt;h2 id="if-you-run-one"&gt;If you run one&lt;/h2&gt;&lt;p&gt;A few practical notes, based on four or five successful runs and some reflection since:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Collect allocations independently before sharing them. Anchoring is real: once the first big bet lands, everyone else's distribution shifts toward it.&lt;/li&gt;
&lt;li&gt;Have whoever spent the fewest points kick off the reveal, then build upwards to the folks with the most conviction, because anchoring will happen. Power dynamics matter too: anyone in a position of authority should go absolutely last.&lt;/li&gt;
&lt;li&gt;Sum the allocations publicly, then talk. The pooled distribution tells you where the room stands; the conversation after the reveal is where the decision actually gets made.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The second point gets to the heart of the problem. A qualitative shouting match becomes a set of quantitative statements of belief, and the team can begin listening to why someone has a certain strength of belief. It's the same instinct as &lt;a href="../../../7/23/going-bayesian-automates-data-analysis/"&gt;going Bayesian with your data analysis&lt;/a&gt;: qualitative hand-flagging of opinions, like qualitative hand-flagging of outliers, gives way to something more principled. Next time your team is going in circles, hand out a thousand points of funny money and see what distributions show up. I suspect you'll be surprised!&lt;/p&gt;
</content></entry><entry><title>Learning how to learn anything, together, in the age of AI</title><link href="https://ericmjl.github.io/blog/2026/8/30/learning-how-to-learn-anything-together-in-the-age-of-ai/" rel="alternate"/><updated>2026-08-30T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:4dff8bb4-20d5-3773-80c5-370a1749bc6b</id><content type="html">&lt;p&gt;Dan Chen and I have a tradition: every year at &lt;a href="https://conference.scipy.org"&gt;SciPy&lt;/a&gt;, we find a corner of the venue, grab coffee, and nerd out about teaching, talking through how people learn, how to teach better, and what actually sticks. The past two years, the conversation has actually taken a pretty sharp turn. Students can now get every answer from AI the moment they sit down with an AI tool, so what is a teacher even for? After sitting with that question for a while, we soon realized the flip side was equally interesting: if someone is genuinely motivated to learn something new, how do they point AI at that goal for maximum leverage?&lt;/p&gt;
&lt;p&gt;That flip is why Dan and I are running &lt;a href="https://learn-anything.nonlinearlabs.ai/"&gt;the Learn Anything with AI retreat&lt;/a&gt;: a five day, in person retreat from the 14th to the 20th of February 2027, with a small, handpicked cohort of people, in the United States. If you click on the link, you'll see the logistics and the details. I want to use this blog post to explain why I want to do this retreat. Here goes.&lt;/p&gt;
&lt;h2 id="we-opened-pandora-s-box"&gt;We opened Pandora's box&lt;/h2&gt;&lt;p&gt;Here's the assumption we're working from: everything related to the boom in AI that we've seen over the past two, three years, they're here to stay. Pandora's box is open. I actually don't see a realistic path back to a world without AI, short of some catastrophe happening. So what's interesting for us as individuals is, what do we do with it? And given that AI is here for good, how do we thrive with it?&lt;/p&gt;
&lt;p&gt;My answer, and Dan's answer as well, is the entire premise of this retreat: to get really good at learning how to learn anything, and how to use AI as the accelerator in that process. If used wrongly, AI will hand you answers without you learning anything. If used correctly, AI can help you build understanding. Because understanding still needs to be built in your head, and that part is the part that people are quietly skipping when they just ask AI for answers. So the people who learn how to learn, with AI's leverage, are going to compound what they know how to do over an entire decade and more. And the people who just collect answers are going to wonder why nothing's stuck. I want to spend five days with a small group of people to explore how we can use AI for the betterment of ourselves.&lt;/p&gt;
&lt;p&gt;The cost of this really needs to be addressed; earlier this year I paid it. I was pumping out stuff using AI models, and by the end of each of those days my brain was fried. I couldn't remember anything I was doing! I was operating faster than the speed of thought, and so I was building what we would call cognitive debt: the equivalent of just cramming for an exam. You cram the inputs in, you skip the deliberate practice, and after the push, the bill arrives in the form of not remembering anything from what you've learned. The antidote is deliberate learning, where we learn it properly, we retain it, and we teach it. And as Dan and I studied and experimented with how to teach and how to learn, here's what we realized: AI can unlock new tools to help us with our understanding, and that turned out to be the real premise of the retreat.&lt;/p&gt;
&lt;p&gt;Want to see the difference between the two paths? Drag the slider and watch what happens to what you remember.&lt;/p&gt;
&lt;iframe src="cognitive-debt.html" width="100%" scrolling="no" style="border: none;" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;h2 id="the-sea-of-change"&gt;The sea of change&lt;/h2&gt;&lt;p&gt;The clearest way I can show you this premise is through my own history. I used to be a lab scientist. Then, after an advisor change during graduate school, I made the choice: I really wanted to move into computing. I knew I had a knack with computers, I just never had the space and time, and this was the perfect space and time to go and make that change for myself. So I flipped the switch on computing. I picked up Python on my own, picked up machine learning on my own, with a little help from some classes, picked up computational biology and the general methods, and soon taught myself network science through books and others' work. I also learned Bayesian statistics in graduate school, self-taught. Throughout that process, something clicked for me: my ability to learn those topics was greatly facilitated by knowing how to do programming. Once I understood programming, I could do math. It wasn't the other way around.&lt;/p&gt;
&lt;p&gt;Most other people, most CS folks, probably get good at math first, and then programming comes as a natural consequence. For me it was the other way around, and I had to learn that about myself. So my ability to learn anything came partly from computing arriving in my life as a tool. But even then, I had some limitations. I couldn't build interactive visualizations for myself. I just didn't understand that paradigm, and building them took too much effort for stuff I wanted to learn, so I just didn't do it and had to find other ways around that limitation. Now with AI, I can have an HTML canvas and build these beautiful visuals that help me understand what's going on, in a Marimo notebook or inside an HTML page. AI can do that heavy lifting for me now. Computing was my first great learning accelerator, and AI is the second. That's the sea of change worth shouting about.&lt;/p&gt;
&lt;p&gt;Here's what that first accelerator looked like up close. &lt;a href="https://www.cs.toronto.edu/~duvenaud/"&gt;David Duvenaud&lt;/a&gt; taught me the backpropagation algorithm in 2016, and how it was really just the chain rule. I'd learned the chain rule in high school, but it didn't really stick for me until graduate school, where the chain rule could be written in Python with NumPy. And once you have the machinery in the form of &lt;a href="https://github.com/HIPS/autograd"&gt;autograd&lt;/a&gt;, you have a way of differentiating through almost any mathematical program you can think of and write. That was really fun. I had to re-derive backprop for myself, starting from linear regression, to really understand how it extends to a larger neural net model very naturally, just by chaining up more functions. That's all a neural net was: chained-up functions. And it was very tedious to get there. But nothing really stuck until I deployed my knowledge in the form of teaching.&lt;/p&gt;
&lt;p&gt;When I committed to teach &lt;a href="https://ericmjl.github.io/teaching/network-analysis-made-simple/"&gt;network analysis&lt;/a&gt; as a tutorial at the scientific Python conference, that commitment was the greatest accelerator of my learning journey. There was one crazy year where I did three tutorials, &lt;a href="https://ericmjl.github.io/teaching/deep-learning-fundamentals/"&gt;deep learning fundamentals&lt;/a&gt;, &lt;a href="https://ericmjl.github.io/teaching/bayesian-data-science-probabilistic-programming-scipy-2019/"&gt;Bayesian statistics&lt;/a&gt;, and network analysis, and I had to solo two of the three because my co-presenters couldn't make it. It was chaotic, and it was a blast. The moment I agreed to teach each of those subjects, I found exactly which parts I had just been hand-waving and never pushed on with real curiosity. It's wonderfully uncomfortable, and it works! That's why the retreat ends with what we call a festival of teaching: everybody teaches what they've learned to the rest of the room, to get that first real rep in.&lt;/p&gt;
&lt;h2 id="in-person-on-purpose"&gt;In person, on purpose&lt;/h2&gt;&lt;p&gt;Beyond the cognitive cost, there's a social one: interacting with AI can be incredibly isolating, because you're sitting alone with a screen, asking questions, and not really interacting with another human. But the whole point of a learning community is other people. One guy I talked with about this retreat put it plainly: he lives in a tech hub, and whenever he goes to networking events, there's always a pitch. There's just no one to interact with in a normal way. So this room we're building, this group of people we're inviting, is essentially a space where you can nerd out with other people, and nobody's selling anything.&lt;/p&gt;
&lt;p&gt;The format is intentionally human centered. We all physically stay together in a retreat center. There are lectures about learning theory, good pedagogy, and how people make things stick, but there's also enough space to duck out when you're peopled out; I'm an introvert, so these escape hatches matter. We're keeping the cohort really small, and we want it to hold a real diversity of knowledge fields. And sharing meals together, five days of shared hallways, is how hallway conversations compound into friendships, just like the SciPy hallway track did for me with Dan Chen. If we run this with overlapping groups, I think this community could grow into one of the more valuable things that any one of us has.&lt;/p&gt;
&lt;h2 id="the-shape-of-the-week"&gt;The shape of the week&lt;/h2&gt;&lt;p&gt;Curious how the five days actually unfold? Step through the week below.&lt;/p&gt;
&lt;iframe src="retreat-week.html" width="100%" scrolling="no" style="border: none;" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Days one through three are the learning science days, and they're Dan's home turf. I've known &lt;a href="https://chendaniely.github.io/"&gt;Dan Chen&lt;/a&gt; for years through the SciPy community; being a professional educator, he has read widely about the science of learning. He's taught multiple classrooms of data science master's students at the University of British Columbia in Vancouver, and he's deeply attuned to how people actually learn and retain stuff, the stuff we were never taught in grad school, or even undergrad. Interspersed throughout those days are sessions to put into practice what you're learning, and practical guidance on where AI slots into each step. As a small example: if you need to generate flashcards for retrieval practice, AI can help you with that. You can even get it to build adaptive systems for your own learning. We'll show you ways to build tools like these for yourself. Everyone arrives having picked a field that's genuinely foreign to them, and starts applying what the days teach to that field, deliberately.&lt;/p&gt;
&lt;p&gt;Day four is more unstructured by design. This is where you home in on prepping to teach the topic you've spent the week learning: you take everything you've absorbed and build the teaching material you'll use to teach someone else the concepts, with plenty of space for creative freedom. And throughout the day, Dan and I will check in with each person once or twice, to see how things are developing and where you're stuck, acting as coaches to push and challenge you toward the day five festival of teaching.&lt;/p&gt;
&lt;p&gt;Day five is the festival of teaching: twenty-some first reps, all in one day. This whole week is a microcosm of what I had to do when I was learning my way through new fields, and here's the goal of it: the retreat is about acquiring the skills needed to use AI to master a brand new field, and we'll spend that week acquiring them, practiced deliberately on a field of your choosing. Stack those skills consistently, month after month, even if you've got kids and a day job, and your capacity to learn compounds in a way that will massively accelerate everything else you take on.&lt;/p&gt;
&lt;p&gt;One more source of inspiration for the retreat design: &lt;a href="https://whonothow.com/"&gt;the book &lt;em&gt;Who Not How&lt;/em&gt;&lt;/a&gt;. We always encourage you to find a who in your life who can act as the subject matter expert for whatever field you're learning, someone you can tap when you have questions, and who you trust to be the ultimate arbiter of whether you've really learned that field. You do want to approach that person with respect for their time. And AI can help you level up your knowledge to the point where you're engaging with that person as a peer, or close to it, rather than in a coach-mentee relationship.&lt;/p&gt;
&lt;h2 id="why-i-keep-calling-it-a-learning-retreat"&gt;Why I keep calling it a learning retreat&lt;/h2&gt;&lt;p&gt;So really, this isn't an AI retreat, and it's not one of those feel-the-spirit kind of retreats either. It's a very practical retreat. The science of learning is the main content. And the main motivation is to bring people together: to build a group of people who are genuinely interested in improving themselves, as I've been able to taste over the past decade of teaching myself multiple things. That's why I keep calling it a learning retreat: the learning comes first, and AI accelerates it.&lt;/p&gt;
&lt;p&gt;If this resonates with you, &lt;a href="https://learn-anything.nonlinearlabs.ai/"&gt;come take a look at the retreat&lt;/a&gt;. We're keeping it small on purpose and making it exclusive: not to a particular demographic, but to a type of personality. Dan and I really hope that we can see you there in February. I think the friends you make in that retreat might be the part that compounds the longest.&lt;/p&gt;
&lt;p&gt;And I'd really love for you to take a small bet on yourself.&lt;/p&gt;
&lt;p&gt;Before you go: five quick statements, and you be the judge of whether they sound like you.&lt;/p&gt;
&lt;iframe src="does-this-resonate.html" width="100%" scrolling="no" style="border: none;" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;</content></entry><entry><title>Quantum Computing for the Probabilistic Bayesian</title><link href="https://ericmjl.github.io/blog/2026/8/7/quantum-ml-for-the-probabilistic-bayesian/" rel="alternate"/><updated>2026-08-07T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:932978ed-018b-3385-ad2c-e4b4cb33937d</id><content type="html">&lt;p&gt;This is a post I have been wanting to write for a while. As of 2026, I have been working much more closely with my teammate Alexey Galda, and part of that has been one-on-one tutoring on the foundations of quantum computing. (To be able to learn from someone well-trained, in a one-on-one setting, is a real privilege!)&lt;/p&gt;
&lt;p&gt;Because I learn best by teaching, and because retrieval practice is the best way to make knowledge stick, I decided to write up what I have absorbed so far. I will readily admit that plenty of details are still beyond my grasp. Even so, I want to leave you with enough of a "&lt;a href="https://basecamp.com/shapeup/1.3-chapter-04#fat-marker-sketches"&gt;fat marker sketch&lt;/a&gt;" of the ideas to reason about where quantum computing might genuinely help.&lt;/p&gt;
&lt;p&gt;Here is the angle I am going to take. I have been doing Bayesian modeling for years, which means I have spent a lot of time thinking about probability distributions, sampling, and inference over spaces too big to enumerate. That turns out to be a surprisingly good lens for quantum computing. So that is how I am going to explain it: quantum mechanics as seen by a probabilistic Bayesian.&lt;/p&gt;
&lt;p&gt;If you're ready for the ride, here we go!&lt;/p&gt;
&lt;iframe src="before-you-read.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;h2 id="the-fundamentals"&gt;The fundamentals&lt;/h2&gt;&lt;p&gt;Start with the most fundamental object in quantum computing: the state of a single qubit. A one-qubit state is an arrow of length 1 in a two-dimensional space whose axes are the two basic states |0&amp;gt; and |1&amp;gt;. The arrow's coordinates along those axes are its amplitudes, which in general are complex numbers. The full one-qubit state space is the Bloch sphere, a unit sphere; when both amplitudes are real, the arrow's tip sits on a unit circle, one great circle of that sphere. That is the slice this post lives on: the signs you can see on it are the drawable case of the complex phases I am leaving out.&lt;/p&gt;
&lt;iframe src="bloch-circle.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Drag the arrow above. Its projections onto the |0&amp;gt; and |1&amp;gt; axes are the amplitudes, and the squared magnitudes of those projections, |amplitude|², are the probabilities of measuring |0&amp;gt; or |1&amp;gt;. That squaring is the bridge into your world as a probabilistic Bayesian, and it is the lens we will use for the rest of the post. The new ingredient, with no analogue in classical probability, is that amplitudes carry phase, which on our circle is just a sign: the arrow can point anywhere on the circle, including where a coordinate is negative. That phase stays quiet for a single measurement, but it is the seed of everything quantum that follows.&lt;/p&gt;
&lt;h3 id="many-qubits-exponentially-many-amplitudes"&gt;Many qubits, exponentially many amplitudes&lt;/h3&gt;&lt;p&gt;That picture was for one qubit, which lives in two dimensions. Add a second qubit and the state now has four basis states (|00&amp;gt;, |01&amp;gt;, |10&amp;gt;, |11&amp;gt;), so the arrow lives in four dimensions, with one amplitude per basis state. Add a third and you get eight. The number of amplitudes doubles with every qubit:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Number of qubits&lt;/th&gt;
&lt;th&gt;Basis states (= amplitudes)&lt;/th&gt;
&lt;th&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0, 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;00, 01, 10, 11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;000, 001, 010, 011, 100, 101, 110, 111&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;n&lt;/td&gt;
&lt;td&gt;2^n&lt;/td&gt;
&lt;td&gt;All binary strings of length n&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Play with the doubling below: one qubit gives two amplitudes, two qubits give four, three give eight.&lt;/p&gt;
&lt;iframe src="many-qubits.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;This is the first place quantum computing leaves classical intuition behind, and it leaves fast. A 50-qubit register has 2^50 amplitudes, more than a quadrillion; you cannot write them all down, let alone store the state on a classical machine. Measuring collapses that whole arrow to a single bitstring, drawn with probability equal to its squared magnitude, the many-qubit version of what the unit circle showed.&lt;/p&gt;
&lt;p&gt;We should clarify one assumption: in the interactive diagrams above, we treated the qubits as independent. The general n-qubit state is a single arrow in 2^n dimensions, and most such arrows cannot be split into one circle per qubit. That non-splitting is entanglement, which has its own section below.&lt;/p&gt;
&lt;p&gt;Before we move on, we should make this picture second nature. Drag the circles until it is obvious how each independent qubit maps to one circle, how the joint amplitudes are the products of their coordinates, and how the probabilities factor. That clean correspondence is the foothold we will want when entanglement arrives and the one-circle-per-qubit picture stops working.&lt;/p&gt;
&lt;h3 id="the-probability-view-coins"&gt;The probability view (coins)&lt;/h3&gt;&lt;p&gt;Step back from amplitudes for a moment. If you only care about what a measurement will tell you, a qubit behaves like a biased coin, and a register of qubits like a pile of coins. That is the probabilistic-Bayesian lens I promised in the intro, and I will lean on it for the rest of the post; the widget below lets you flip 1, 2, or 3 and watch the outcomes follow the distribution. Just remember the coin view is a simplification, the shadow the amplitudes cast on a measurement, and it hides the signs that the interference section later puts to work.&lt;/p&gt;
&lt;iframe src="qubit-measurement.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;h3 id="entanglement"&gt;Entanglement&lt;/h3&gt;&lt;p&gt;Entanglement, in this framing, is when the state of one qubit cannot be described independently of another. When two qubits are entangled, their measurement outcomes are correlated, no matter how far apart they are. No information travels between them: each outcome is random on its own, and the correlation only shows up when you compare results over an ordinary classical channel. This non-local correlation is a key resource in quantum computing, and it has no clean classical analogue. For a probabilistic Bayesian, think of it as two coins whose flips look random individually but line up once you compare records.&lt;/p&gt;
&lt;p&gt;Here is that idea made geometric. Two dials below set the qubits as if they were independent; the entanglement dial mixes in correlation. Watch the joint amplitudes stop factorizing, and the concurrence climb from 0 (independent) toward 1 (maximally entangled).&lt;/p&gt;
&lt;iframe src="two-qubit-entanglement.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;h3 id="interference"&gt;Interference&lt;/h3&gt;&lt;p&gt;An amplitude's sign is easy to ignore, because a single measurement only sees the squared size, which is just a probability. The sign only matters when amplitudes combine: two amplitudes of the same sign reinforce, and two of opposite sign can subtract all the way to zero. That cancellation is &lt;em&gt;interference&lt;/em&gt;, and it has no classical analogue. Classical probability distributions are built from non-negative numbers that pile up; quantum states are built from signed amplitudes that can interfere.&lt;/p&gt;
&lt;p&gt;One tool that does the mixing is the Hadamard gate. It takes any input state and produces a specific output by blending the two input amplitudes together, with a minus sign baked into one of the blends. The widget below lets you drag the input arrow and watch the output respond. The butterfly diagram in the middle traces each input amplitude as it splits into contributions at each output, so you can see exactly where reinforcement and cancellation happen.&lt;/p&gt;
&lt;iframe src="hadamard-gate.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Drag the input to the diagonal (45 degrees, the |+&amp;gt; state) and look at the butterfly. Two contributions arrive at each output. At output |0&amp;gt;, both are positive, so they reinforce, and the output arrow swings hard toward |0&amp;gt;. At output |1&amp;gt;, one is positive and one is negative, so they cancel, and |1&amp;gt; vanishes. That is interference made visible: the minus sign killed one output while boosting the other. Now drag to other angles and watch how the balance between reinforcement and cancellation shifts. When only one input amplitude is nonzero (try |0&amp;gt; or |1&amp;gt;), there is only one path through the gate, so nothing cancels.&lt;/p&gt;
&lt;p&gt;Two Hadamards in a row undo each other precisely because the minus signs cancel on the way back. On &lt;em&gt;n&lt;/em&gt; qubits you apply a Hadamard to each wire; that is how a uniform superposition gets built in the first place, and how it gets recombined later so the signs can interfere.&lt;/p&gt;
&lt;h3 id="from-interference-to-answers"&gt;From interference to answers&lt;/h3&gt;&lt;p&gt;A quantum program is an arrangement of gates that concentrates amplitude on the qubit configurations representing the solution to a problem, and cancels the rest. The gates do not need to know the answer ahead of time; they encode a way to score candidates, and interference amplifies whatever scores well. The result is a probability distribution deliberately shaped so that the solution is likely and the noise has vanished.&lt;/p&gt;
&lt;p&gt;If you are a Bayesian, this should sound familiar. Your prior and likelihood shape a posterior; a quantum circuit's gates shape amplitudes. Then, to read the answer out, you measure, and measurement is just drawing a sample from that shaped distribution.&lt;/p&gt;
&lt;p&gt;And here is where it finally clicks for me, because in my own work building Bayesian models, sampling is the bottleneck. The core problem in Bayesian inference never changes: your posterior may be intractable, so you sample. But sampling is hard! Your chains may not mix, or there may be divergences, and if part of your problem lives in a combinatorial space, good luck with the approximations.&lt;/p&gt;
&lt;p&gt;What if you didn't have to enumerate a combinatorial space at all, but instead let your sampler live in superposition over the whole of it?&lt;/p&gt;
&lt;p&gt;That's the quantum computer's party trick. It isn't a magical exponential speedup for everything. But calling it an MCMC sampler goes too far; generic circuits evolve amplitudes coherently rather than stepping through a Markov chain, and true quantum MCMC algorithms are special cases rather than the norm. View it instead as a sampler purpose-built for exactly the kind of problem we already wrestle with, inference over spaces too big to enumerate classically.&lt;/p&gt;
&lt;iframe src="check-understanding.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;h2 id="the-takeaway"&gt;The takeaway&lt;/h2&gt;&lt;p&gt;In this blog post, we stripped away much of the physics of quantum computing to give you the following picture to walk away with. A qubit is an arrow on a unit circle, the real slice of the Bloch sphere; its coordinates are amplitudes, and their squared magnitudes are probabilities. A register of &lt;em&gt;n&lt;/em&gt; qubits is a single arrow in 2^n dimensions, one amplitude per possible bitstring. Entanglement is when that arrow cannot be split into one circle per qubit. Interference is what happens when signed amplitudes combine through gates: same signs reinforce, opposite signs can cancel, and the quantum program uses this to concentrate amplitude on the configurations that represent solutions. Measurement is drawing a sample from the resulting distribution.&lt;/p&gt;
&lt;p&gt;If you are a probabilistic Bayesian, the parallel is exact. Your prior and likelihood shape a posterior; a quantum circuit's gates shape amplitudes. You sample from your posterior; you measure from the quantum state. The quantum computer is not a faster laptop or a magic box. It is a device that is very good at exactly the thing we already struggle with classically: representing and sampling from distributions over spaces too big to enumerate.&lt;/p&gt;
&lt;p&gt;Many details here are still beyond my grasp. But holding this mental model has already changed how I read about quantum computing. Where I used to see mysterious physics, I now see amplitude vectors, clever encodings, and interference doing the heavy lifting, all on hardware built for exactly that. I hope it does the same for you.&lt;/p&gt;
</content></entry><entry><title>How AI turbocharges your learning</title><link href="https://ericmjl.github.io/blog/2026/7/27/ai-turbocharges-learning/" rel="alternate"/><updated>2026-07-27T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:23f849b9-5e51-31f2-a7d9-2e7cd8767c2a</id><content type="html">&lt;p&gt;I've been sitting on a question since SciPy this year. Daniel Chen and I get together every conference to talk about education, it's our shared obsession, and this time the big one was: AI shortcuts everything now. How do we show people the patterns for learning deeply with it rather than letting it replace the learning? I didn't have a good answer. So I went and read the research. This post is what I found.&lt;/p&gt;
&lt;p&gt;The finding, repeated in every paper and every story: AI doesn't help or harm learning on its own. &lt;strong&gt;It amplifies whatever cognitive habits you bring to it.&lt;/strong&gt; Bring good study habits and the tool multiplies them. Bring passivity and it multiplies that too.&lt;/p&gt;
&lt;p&gt;None of the techniques are new. Retrieval practice, spaced repetition, elaboration, the Socratic method, they've been around for decades. What's new is that AI can finally run all of them for one learner at a time, at a scale no human tutor can match.&lt;/p&gt;
&lt;h2 id="before-you-read"&gt;Before you read&lt;/h2&gt;&lt;p&gt;Before you read on, I'd like to invite you to try these three, even if you're just guessing. Recalling first is what makes the rest of the post stick. Then keep reading. By the end, you'll know which of your answers the evidence supports.&lt;/p&gt;
&lt;iframe src="before-you-read.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;h2 id="the-central-distinction-amplifier-or-substitute"&gt;The central distinction: amplifier or substitute&lt;/h2&gt;&lt;p&gt;The cleanest version of the thesis comes from a 2025 study out of Turkey, run by &lt;a href="https://hamsabastani.github.io/education_llm.pdf"&gt;Bastani and colleagues&lt;/a&gt; and published in PNAS. They split roughly a thousand students learning math into two groups. Both groups used GPT-4, with the same content and the same kids. The only thing that changed was the prompt.&lt;/p&gt;
&lt;p&gt;The first group got plain ChatGPT, asked questions, and copied answers. It felt productive during practice. But when the AI was taken away for the final exam, they scored 17 percent worse than students who never had AI at all. They had used it as a crutch and never built the skill. The second group got a hinted, grounded tutor instead, one that withheld answers and asked leading questions. During practice they performed 127 percent better, and when the AI was taken away, they held their ground.&lt;/p&gt;
&lt;p&gt;The authors call this the Dual-Mechanism Model. AI can amplify your thinking when the interaction forces you to do it, or substitute for it when you hand it over. &lt;strong&gt;The same tool, used two ways, produces opposite results, and the difference is always whether the AI does the thinking or you do.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Here's an example of that distinction. A student two weeks before her organic chemistry final types, "Explain SN1 and SN2 reactions to me." The model produces a beautiful explanation. She reads it, feels that she understands the material, but actually doesn't. The problem? Familiarity is not the same as understanding, and similarly, recognition is not retrieval. Two weeks later, walking into the exam, she can't tell the two mechanisms apart under pressure.&lt;/p&gt;
&lt;p&gt;Flip the prompt and everything changes. Instead of "explain X to me," she types: "Here is what I think distinguishes SN1 from SN2. SN1 has a carbocation intermediate and unimolecular rate-determining step. SN2 is a concerted backside attack. What am I missing?" Now she's retrieving. The model can correct her, surface the edge cases she didn't mention (solvent effects, stereochemistry, rearrangements), and quiz her on them. She's doing the thinking, and the model is the infinitely patient examiner. That's the whole game in one prompt swap!&lt;/p&gt;
&lt;p&gt;To see if you got the idea, try the little quiz below. Decide for each whether the learner is amplifying their thinking or substituting for it.&lt;/p&gt;
&lt;iframe src="prompt-inspector.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;If you noticed the pattern, there's really only one question that matters: who's generating the thinking, and thus the words? Every study and every story below comes back to that.&lt;/p&gt;
&lt;h2 id="what-the-evidence-actually-says"&gt;What the evidence actually says&lt;/h2&gt;&lt;p&gt;But that's one study. The harder question is whether the whole literature agrees, and on that we finally have meta-analyses. Four large reviews, published in 2025 and 2026, together covering well over a hundred studies and tens of thousands of learners.&lt;/p&gt;
&lt;p&gt;On average, AI helps learning. Across &lt;a href="https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1758670/full"&gt;36 controlled studies and 7,229 participants&lt;/a&gt;, the effect size is g = 0.499, which sounds modest until you realize most education interventions hover near zero. This is medium-to-large.&lt;/p&gt;
&lt;p&gt;But the average is hiding the interesting part. When the researchers dug into what predicted the biggest gains, one variable dwarfed everything else: how the AI was used. Collaborative, structured, Socratic use hit g = 1.026. Traditional, unguided use limped to g = 0.470. &lt;strong&gt;The teaching method wrapped around the AI is what produces that 2x gap.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A &lt;a href="https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1863931/full"&gt;second meta-analysis&lt;/a&gt;, this one covering 89 studies, found that 40.4 percent of AI deployments helped and 23.6 percent hurt, with the rest mixed. The leading risk factor, named in 33.7 percent of studies, was over-reliance. And the most damning number in the whole review: 55.1 percent of studies used no specified pedagogical strategy at all. Much of what gets reported as "AI harms learning" is really "unstructured AI use harms learning."&lt;/p&gt;
&lt;p&gt;To help you see the pattern for yourself, I collated the studies into the map below. Click any one for the design, the sample, and the effect size. Here's what to observe: scaffold the AI and the effects are large; leave it unguided and they flip.&lt;/p&gt;
&lt;iframe src="evidence-map.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Two studies in that map are worth slowing down for. One shows what a good scaffold buys you; the other shows what happens when you don't use one. At the strong end, a &lt;a href="https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf"&gt;World Bank trial in Nigeria&lt;/a&gt; gave roughly 800 students free Microsoft Copilot alongside a six-week afterschool program, with tutors grounding the AI in the curriculum. The result was a 0.31 standard deviation gain, which the authors translate into roughly 1.5 to 2 years of business-as-usual schooling in Nigeria compressed into six weeks.&lt;/p&gt;
&lt;p&gt;At the other end, the &lt;a href="https://arxiv.org/abs/2506.08872"&gt;MIT EEG study&lt;/a&gt; gave students ChatGPT with no scaffold, no tutor, no curriculum. 83 percent couldn't correctly quote their own essay afterward, and brain connectivity was the weakest of any group in the study.&lt;/p&gt;
&lt;h2 id="the-cost-of-cognitive-surrender"&gt;The cost of cognitive surrender&lt;/h2&gt;&lt;p&gt;There's a name for what happens when you lean on the model a little too hard. Ethan Mollick calls it cognitive surrender. You stop thinking, the model drives, and because words are still hitting the page you feel productive. (You did produce a document, so the feeling isn't entirely lying to you.) The learning, though, left the room.&lt;/p&gt;
&lt;p&gt;The MIT study is the one everybody reaches for when they want to say AI rots your brain. I sat with it for an afternoon, and my honest read is that the headline is stronger than the evidence. The brain-connectivity finding? Eighteen people in the final session. The "growing dependence" signal? Could just be practice. But the quote-your-own-essay finding is rock solid: 83 percent of ChatGPT users couldn't accurately reproduce the essay they'd just written. You outsource the writing, you don't even encode it. Hold on to that number.&lt;/p&gt;
&lt;p&gt;There is a longer history here. In 2011, &lt;a href="https://doi.org/10.1126/science.1207745"&gt;Sparrow and colleagues at Harvard&lt;/a&gt; found that people who expect to be able to look something up later remember it less well, a finding they called the Google effect. ChatGPT is Google with better prose. The cognitive offloading is the same, only smoother, which makes it more dangerous because it doesn't feel like offloading. It feels like thinking.&lt;/p&gt;
&lt;p&gt;A &lt;a href="https://doi.org/10.17605/osf.io/xk7ta"&gt;2026 systematic review of 44 studies on cognitive offloading&lt;/a&gt; names the mechanism precisely. Offloading lower-order work frees capacity, but the benefit only converts into deeper learning under two conditions: the learner must reallocate the freed capacity to harder reasoning, and the learner must have enough metacognitive awareness to notice that the capacity was freed in the first place. &lt;strong&gt;Without those two conditions, AI is a crutch. With them, it is a lever.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="the-mechanics-retrieval-spacing-and-desirable-difficulty"&gt;The mechanics: retrieval, spacing, and desirable difficulty&lt;/h2&gt;&lt;p&gt;Three principles make AI tutoring work: retrieval practice, spacing, and the testing effect. None are new; all are decades old. What's new is that AI can finally run them for every learner at once.&lt;/p&gt;
&lt;p&gt;The first principle, retrieval practice, has some of the strongest evidence in all of learning science. &lt;a href="https://doi.org/10.1111/j.1467-9280.2006.01693.x"&gt;Roediger and Karpicke&lt;/a&gt; ran the classic 2006 experiment: students who took a brief test after reading recalled far more later than students who reread, even though the rereaders felt more confident. That finding replicates everywhere, from word lists to medical school curricula. Looking away and recalling produces retention. Rereading produces familiarity.&lt;/p&gt;
&lt;p&gt;The second principle, spacing, is counterintuitive: letting a memory almost fade before you review it makes it stronger, not weaker. Ebbinghaus mapped the forgetting curve back in 1885, and the shape is steep. Memory drops fast after a study session. But if you reactivate the memory just before it would have slipped away, the effort of retrieving it builds a far stronger trace than rereading at full strength. That effort is what psychologists call desirable difficulty.&lt;/p&gt;
&lt;p&gt;The third principle is the testing effect, and it's what happens when retrieval practice compounds. Each additional retrieval, especially under slight variations of the question, builds the memory further, in a way a single exposure never could. Every medical school Anki deck runs on this principle.&lt;/p&gt;
&lt;p&gt;Want to see the forgetting curve in action? Pick a strategy below and watch what happens to retention over two weeks.&lt;/p&gt;
&lt;iframe src="forgetting-curve.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;The &lt;a href="https://doi.org/10.48550/arxiv.2309.13060"&gt;UniDistance study&lt;/a&gt; is the cleanest field test of all three principles at once. Researchers gave 51 students a semester-long AI tutor (the MAGMA Learning app) that used a neural network to model each student's grasp of every concept, then served questions at each student's moment of near-forgetting, and visualized mastery as a 3D "learnet" whose brightness tracked grasp. Spacing, retrieval, and personalization, all automated.&lt;/p&gt;
&lt;p&gt;The students gained up to 15 percentile points over the semester, and the model's predicted grasp correlated with actual exam grade at r = 0.81. The boring learning science works! AI can finally run all of it for each learner at once.&lt;/p&gt;
&lt;h2 id="the-prompt-swap-and-its-cousins"&gt;The prompt swap and its cousins&lt;/h2&gt;&lt;p&gt;Up to this point, the focus has been on what apps and tutors do for you. Let's flip it: what can you do for yourself, with whatever model you're already using? Pick ChatGPT, Claude, or Gemini; it doesn't matter. How you talk to the model is what counts, and that skill outlasts any specific app.&lt;/p&gt;
&lt;p&gt;Start with the prompt swap. It's the single highest-leverage move I know. Take any prompt where you'd have typed "explain X to me" and rewrite it as "here is what I think X is, am I missing anything?" You've just flipped the model from answer machine to coach, and forced yourself to retrieve instead of recognize.&lt;/p&gt;
&lt;p&gt;You can push this further with a Socratic prompt. Tell the model: "Don't give me the answer. Ask me questions, one at a time, that lead me to discover it." The Socratic method is centuries old, but AI can finally run it for free for anyone with a phone. A &lt;a href="https://arxiv.org/pdf/2508.05116"&gt;2025 Nature Human Behaviour paper&lt;/a&gt; found that this approach won by asking better questions, which activated what the authors call epistemic agency: the learner's ownership of their own reasoning.&lt;/p&gt;
&lt;p&gt;Here's another one. Flip the roles: ask the model to be the student, not the teacher. "I'm going to teach you X as if you're a beginner. Stop me whenever I say something unclear or wrong." You can't teach what you don't own, and the model is a tireless beginner who won't let you wave your hands.&lt;/p&gt;
&lt;p&gt;The ACTOR framework, from Sandeep Swadia, systematizes this for reading. Aim (state in one sentence why you are reading this). Compress (name the trunk of the idea tree before the leaves). Test (read to reject, not to agree; ask the model to challenge your interpretation and find the hidden assumption). Own (restate in your own words; connect to a real meeting, mistake, or person; teach it). Run (convert the idea into one decision, one rule, one checklist, one experiment). A book should interrupt your behavior, not just your beliefs.&lt;/p&gt;
&lt;p&gt;These four moves are the toolkit. Next, let's see what happens when real people put them to work.&lt;/p&gt;
&lt;h2 id="in-practice-stories-from-people-who-did-it"&gt;In practice: stories from people who did it&lt;/h2&gt;&lt;p&gt;&lt;a href="https://dev.to/tomerl1/i-know-rust-1jpi"&gt;Tomerl1&lt;/a&gt; set out to learn Rust in eighteen days by building a game, with Claude as a mentor and a hard rule: zero AI-written production code. He could paste code into the model, but he couldn't paste the model's code into his project. His summary, half a year later: "Any bug I fixed by pasting into Claude and pasting the answer back came back as a different bug a week later. The fixes that stuck were the ones I understood."&lt;/p&gt;
&lt;p&gt;Next is &lt;a href="https://medium.com/@katy.saintin/how-i-got-ai-certified-using-an-ai-that-had-no-teaching-degree-0da2a3575d43"&gt;Katy Saintin&lt;/a&gt;, an industrial software engineer who ran a two-window experiment. She opened Claude twice, side by side: one mentor, one shortcut. Only the mentor window taught her anything. But what stuck with me was her observation about her colleagues. In her world (SCADA, EPICS, TANGO control systems), junior engineers stay quiet because asking feels embarrassing. "The AI never laughed at a beginner question. Never rolled its eyes. Never made curiosity feel expensive." The real disruption might be simpler than we think: for the first time, millions of people feel safe enough to ask questions.&lt;/p&gt;
&lt;p&gt;AI removes the friction of making practice exams, but you still have to take them. That's the insight &lt;a href="https://medium.com/@JordanJPiece/how-to-use-ai-to-ace-your-exams-aa6141fdf9da"&gt;Jordan Pierce&lt;/a&gt; took from scoring 96 on a finance exam whose average was 74, then 100 out of 100 twice. He had the model write neighbor-confusion distractors (wrong answers that look right), match the question length, build novel scenarios, and repeat any error until he'd fixed it three times in a row. The AI was a test-writing assistant. He was the one taking the tests.&lt;/p&gt;
&lt;h2 id="the-meta-skill"&gt;The meta-skill&lt;/h2&gt;&lt;p&gt;What's the single biggest predictor of how much AI helps you learn? It's you.&lt;/p&gt;
&lt;p&gt;A &lt;a href="https://www.mdpi.com/2079-3200/13/12/160"&gt;2025 meta-analysis in Educational Psychology Review&lt;/a&gt;, covering 29 experiments, found that the single largest moderator was the learner's prior self-regulation. High self-regulated learners gained g = 0.863; low self-regulated learners gained g = 0.284. What I want you to notice is the size of that gap. It's larger than the average effect of AI itself.&lt;/p&gt;
&lt;p&gt;AI amplifies what you bring. If you already know how to study, it supercharges your study. If you don't, it mostly supercharges your ability to produce documents that look like studying.&lt;/p&gt;
&lt;p&gt;Bloom's lower levels (remember, understand, apply) are commoditized. The model does them faster than you. Value has moved up to analyze, evaluate, create. &lt;strong&gt;The differentiator is what you do on top of having the information.&lt;/strong&gt; The meta-skill that compounds is driving higher-order questioning and judgment.&lt;/p&gt;
&lt;iframe src="blooms-taxonomy.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;So what's the counter to "brain rot"? Self-directed learning. The people who thrive with AI are the ones who drive the questioning the model won't generate on its own, check the output against their own understanding, and treat the model as a coach rather than an oracle. It's a habit of mind, not a feature you can toggle.&lt;/p&gt;
&lt;h2 id="check-your-understanding"&gt;Check your understanding&lt;/h2&gt;&lt;p&gt;Those three questions I invited you to try at the top? You now have the evidence to answer every one. Here's a five-question quiz to check what stuck.&lt;/p&gt;
&lt;iframe src="check-understanding.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Here's my parting advice: take the prompt swap. Next time you open a model to learn, tell it what you think and ask where you're wrong. The model won't be any different. You will.&lt;/p&gt;
</content></entry><entry><title>The understanding zeitgeist in AI code review</title><link href="https://ericmjl.github.io/blog/2026/7/24/understanding-zeitgeist-ai-code-review/" rel="alternate"/><updated>2026-07-24T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:a65682c4-d37b-3dd0-bb3f-f5d685987172</id><content type="html">&lt;p&gt;I've spent the year on both ends of agent-written pull requests. I've opened them after a long coding session, half-hoping the reviewer catches what I only half-tracked. And I've reviewed them, eyes glazing over a forty-file diff, typing "looks good to me" while a quieter voice asks whether it actually does. (Reader, I've been the problem as often as the solution.) If you've used a coding agent, you know the feeling.&lt;/p&gt;
&lt;p&gt;This spring and summer, a bunch of people I read arrived at the same worry from completely different angles: now that agents write so much of our code, producing it is cheap, but &lt;em&gt;understanding&lt;/em&gt; it is the expensive part. I want to map where the conversation has landed, so you can see its shape and find your own place in it.&lt;/p&gt;
&lt;p&gt;One warning, and a promise. The worry has a name now, and the people working on it have built real techniques for it. I'm going to steal one of those techniques here, because Geoffrey Litt and Andy Matuschak both argue that reading is a weak way to build understanding; you have to be quizzed on it. So this post practices what it preaches. Three quick questions up front, to surface what you already think, and five at the end, to check what landed. (Yes, I'm quizzing you on a post about why quizzes work. I couldn't resist.)&lt;/p&gt;
&lt;h2 id="before-you-read"&gt;Before you read&lt;/h2&gt;&lt;p&gt;These aren't graded, and the cards won't tell you if you're right. Pick your honest answer for each, notice your assumptions, then read on. The survey revisits every one.&lt;/p&gt;
&lt;iframe src="before-you-read.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Hold your answers. Let's see where the field has landed.&lt;/p&gt;
&lt;h2 id="understanding-is-the-bottleneck"&gt;Understanding is the bottleneck&lt;/h2&gt;&lt;p&gt;Geoffrey Litt gave a talk at the AI Engineer conference in July 2026 with a thesis that names the whole conversation: &lt;strong&gt;understanding is the new bottleneck.&lt;/strong&gt; Writing code got cheap. Holding a mental model of what the code does is the expensive part now, and it's the gap every pull request asks us to cross.&lt;/p&gt;
&lt;p&gt;The deeper root of this worry predates the AI boom by decades. &lt;a href="https://margaretstorey.com/blog/2026/02/09/cognitive-debt/"&gt;Margaret Storey&lt;/a&gt;, reviving Peter Naur's 1985 essay &lt;a href="https://pages.cs.wisc.edu/~remzi/Naur.pdf"&gt;"Programming as Theory Building"&lt;/a&gt;, puts it precisely: a program is a theory that lives in the minds of the developers, capturing what the software does and how it can change. Technical debt lives in the code. &lt;strong&gt;Cognitive debt lives in the developers' minds&lt;/strong&gt;, and Storey argues it will paralyze a team before technical debt does. She watched it happen to a student team that could no longer make simple changes, because nobody could explain why the system worked the way it did.&lt;/p&gt;
&lt;p&gt;This is the through-line of the whole survey. The people I've been reading disagree on tactics but agree on the diagnosis. When typing the code stops being the bottleneck, the bottleneck moves to whoever has to hold the theory.&lt;/p&gt;
&lt;h2 id="what-breaks-at-the-pull-request"&gt;What breaks at the pull request&lt;/h2&gt;&lt;p&gt;The diagnosis bites hardest at the pull request, because that's the moment the theory has to move from author to reviewer. Several practitioners wrote the reviewer's side raw this year, and it hits closest to home.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://hojberg.xyz/the-programmer-identity-crisis/"&gt;Simon Højberg&lt;/a&gt; names the lived experience: code reviewers are "rapidly losing their minds" because they've become "the first layer of quality control instead of one of the last," picking apart hallucinated libraries and uncalled functions while the author shrugs that Claude wrote it. I recognize the feeling; I gloss over the diff, overwhelmed and bored, and approve whatever's there as long as CI is green.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://tim-schipper.nl/en/blog/reviewing-ai-generated-pull-requests"&gt;Tim Schipper&lt;/a&gt; explains why agent PRs are uniquely hard to review, with a framing I've stolen more than once. &lt;strong&gt;An AI PR has no smell.&lt;/strong&gt; It's uniformly polished, because polish is the one thing next-token prediction is genuinely good at, so the reviewer's instinct fires the wrong way: the surface looks more reviewable than a human's, so it gets less review. His test for whether you actually reviewed it is the sharpest line in the whole conversation: approve a PR you couldn't have written and can't fully explain, and you didn't review it. You witnessed it.&lt;/p&gt;
&lt;p&gt;Here is what that looks like in practice: a clean diff, and the five checks a reviewer actually has to run.&lt;/p&gt;
&lt;iframe src="ai-pr-no-smell.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Every checklist I read hits the same failure modes. &lt;a href="https://github.blog/ai-and-ml/generative-ai/agent-pull-requests-are-everywhere-heres-how-to-review-them/"&gt;GitHub's engineering team&lt;/a&gt; published guidance in May 2026 built around the insight that agents optimize for surface plausibility, not semantic correctness. Their data is striking: GitHub Copilot code review has processed over 60 million reviews, more than one in five code reviews on GitHub now involve an agent, and a January 2026 study found agent-generated code introduces more redundancy and quiet debt per change than human-written code. The playbook targets the specific ways agents slip: gaming CI to go green, reimplementing utilities that already exist, hallucinating imports that resolve to the wrong symbol, and writing tests that pass without verifying anything.&lt;/p&gt;
&lt;p&gt;That last one is the most insidious. If the same agent wrote the code and the tests in the same run, the green checkmark is nearly worthless, because a model that made a wrong assumption in the code makes the identical wrong assumption in the test. The strongest heuristic is the same in every checklist: &lt;strong&gt;demand a test that fails on the pre-change behavior.&lt;/strong&gt; If the agent can't write a test that would have caught the bug it claims to fix, the fix is incomplete or the understanding is wrong.&lt;/p&gt;
&lt;h2 id="the-cost-of-velocity-without-understanding"&gt;The cost of velocity without understanding&lt;/h2&gt;&lt;p&gt;Why does this bite? Because agents let us ship changes that we'd normally have deliberated over for weeks in a matter of hours, and the mismatch between output speed and comprehension speed is where the debt compounds.&lt;/p&gt;
&lt;p&gt;Mario Zechner, who built the Pi agent framework, put it bluntly in a post &lt;a href="https://simonwillison.net/2026/Mar/25/thoughts-on-slowing-the-fuck-down/"&gt;Simon Willison&lt;/a&gt; amplified. With an orchestrated army of agents there's no human bottleneck, and the mistakes compound faster than any one person would let them:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;These tiny little harmless booboos suddenly compound at a rate that's unsustainable. You have zero fucking idea what's going on because you delegated all your agency to your agents. You let them run free, and they are merchants of complexity.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Willison's own takeaway is that cognitive debt is real, and that we now need a new balance of speed against mental thoroughness, since typing the code is no longer anywhere close to the bottleneck on writing software.&lt;/p&gt;
&lt;p&gt;The cost shows up as burnout too. Steve Yegge describes the "AI vampire" effect: agents automate the easy work and leave you with all the difficult decisions, which is exhausting in a way raw output isn't, and he finds four hours of agent work a day a more realistic pace than eight. &lt;a href="https://hbr.org/2026/02/ai-doesnt-reduce-work-it-intensifies-it"&gt;A Berkeley Haas study in HBR&lt;/a&gt; found AI intensifies work rather than reducing it, creating continual attention-switching and a sense of always juggling. Velocity without understanding produces brittle code, and it produces tired people.&lt;/p&gt;
&lt;h2 id="understand-to-participate"&gt;Understand to participate&lt;/h2&gt;&lt;p&gt;Litt flips the question here. If agents are getting better at verifying their own work, what's left for the human reviewer?&lt;/p&gt;
&lt;p&gt;The tempting answer is to keep checking the code. But the people I read are converging on a different answer. &lt;strong&gt;We understand to participate.&lt;/strong&gt; A project is hundreds of iterative loops with the agent, and the mental model you carry into the next loop is exactly what lets you come up with the next idea. Litt's phrase, which &lt;a href="https://simonwillison.net/2026/Jul/2/understand-to-participate/"&gt;Willison picked out&lt;/a&gt; as the framing that resonated most, is "understand to participate." Verification is the part agents are eating. Participation is the part that's left to us, and it's the part that matters.&lt;/p&gt;
&lt;p&gt;This reframe has old, deep roots. Litt closes his talk by reaching back to Alan Kay and Douglas Engelbart: the point of computing was always to augment, not just automate. &lt;a href="https://distill.pub/2017/aia/"&gt;Shan Carter and Michael Nielsen&lt;/a&gt; made the distinction sharp in a 2017 Distill essay: cognitive outsourcing hands a task to an oracle, while cognitive transformation invents new representations that expand the range of human thought. A reviewer who only checks correctness is doing cognitive outsourcing. A reviewer who builds a richer mental model is doing cognitive transformation, and that's the goal.&lt;/p&gt;
&lt;h2 id="techniques-borrowed-from-teaching"&gt;Techniques borrowed from teaching&lt;/h2&gt;&lt;p&gt;If understanding is the bottleneck, how do we cross it faster? Litt's move is to stop trying to review harder and steal instead from a field that's worked on this for centuries: education. His talk lays out three techniques, and the first two are things an agent can produce while it reviews.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Explanations.&lt;/strong&gt; Litt's &lt;a href="https://gist.github.com/geoffreylitt/a29df1b5f9865506e8952488eac3d524"&gt;&lt;code&gt;/explain-diff&lt;/code&gt;&lt;/a&gt; skill turns a diff into a structured explainer: background first, then the goal and the intuition behind it, then the code walked through in a sensible order with prose around it. A typical diff hands you forty files in alphabetical order; a literate diff hands you a story you can follow.&lt;/p&gt;
&lt;p&gt;Toggle between the two for the same change:&lt;/p&gt;
&lt;iframe src="literate-diff.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;This is where reviewing big PRs with agents pays off most, since the bigger the change, the more the raw diff fails and the explainer earns its keep. Nested inside every explainer is a second device Litt borrowed from Andy Matuschak: a five-question quiz you answer before you approve. His rule is one I've adopted wholesale; he won't send code to a teammate until he can pass the quiz, and he does the same when he reviews theirs. The agent loop runs faster than the speed of human understanding, and the quiz is the mechanical brake that asks whether you actually get it. (This post is my own attempt at that brake; the widget at the end is the quiz, built from &lt;a href="https://andymatuschak.org/books/"&gt;Matuschak's argument&lt;/a&gt; that "books don't work" and the spaced-repetition pattern he and Nielsen embedded in Quantum Country.)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Micro-worlds.&lt;/strong&gt; This one's my favorite. Simon Willison arrived at the same idea independently and calls them &lt;a href="https://simonwillison.net/guides/agentic-engineering-patterns/interactive-explanations/"&gt;interactive explanations&lt;/a&gt;. Instead of describing what a change does, have the agent build a little environment you can poke. Litt migrated his website between frameworks and couldn't review the agent's script, so he asked the agent to build a command center where he clicked through the migration step by step and watched the two sites side by side. Willison wanted to understand a word-cloud algorithm whose code only told him it used "Archimedean spiral placement," so he had an agent animate the placement, watching each word try a spot, overlap, and spiral outward until the idea clicked.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shared spaces.&lt;/strong&gt; The first two are solo techniques; the third is about understanding &lt;em&gt;together&lt;/em&gt;. Litt's point is that when a team holds the same mental model they can communicate efficiently, jam, and riff, and without those shared structures the creative conversations stall. It loops straight back to Storey and Naur: the theory of the system is distributed across the team, and the review is one of the moments that theory gets rebuilt and shared.&lt;/p&gt;
&lt;p&gt;Here's the thread across all three: &lt;strong&gt;agents can write code that helps you understand other code.&lt;/strong&gt; Nathan Baschez adds a trick that works at the planning stage: ask the agent for two versions of every plan, one highly technical and detailed for itself, and one "entertaining essay designed to build my intuition" for you. The second version exists to hand you the mental model.&lt;/p&gt;
&lt;h2 id="huds-over-copilots"&gt;HUDs over copilots&lt;/h2&gt;&lt;p&gt;These techniques imply something about what good AI review tooling should feel like, and &lt;a href="https://geoffreylitt.com/2025/07/27/enough-ai-copilots-we-need-ai-huds"&gt;Litt makes the distinction explicit&lt;/a&gt; in a separate essay. A tool that returns approve or request changes is a copilot, a virtual assistant you talk to. &lt;strong&gt;A tool that builds a micro-world around the change is a HUD&lt;/strong&gt;, a new sense layered into your field of view, like spellcheck's red squiggles. Litt's argument, drawing on Mark Weiser's 1992 rant against the agent metaphor, is that we have enough copilots. We need HUDs!&lt;/p&gt;
&lt;p&gt;Practitioners arrive at the same place: the human role in review is narrowing from correctness verification to context verification, judging whether the agent understood the codebase, the contract, and the intent. Automated review owns the mechanical first pass, and human attention goes to the semantic judgment agents can't self-check. Copilots automate the verdict. HUDs extend your senses so you can make that judgment yourself.&lt;/p&gt;
&lt;h2 id="the-zeitgeist-in-one-map"&gt;The zeitgeist in one map&lt;/h2&gt;&lt;p&gt;Here's the whole conversation compressed. Each one I read contributes a single move to the argument.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it contributes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://geoffreylitt.com/2026/07/02/understanding-is-the-new-bottleneck"&gt;Geoffrey Litt&lt;/a&gt;, "Understanding is the new bottleneck"&lt;/td&gt;
&lt;td&gt;The diagnosis; the understand-to-participate reframe; explanations, micro-worlds, shared spaces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://margaretstorey.com/blog/2026/02/09/cognitive-debt/"&gt;Margaret Storey&lt;/a&gt; / Peter Naur&lt;/td&gt;
&lt;td&gt;Cognitive debt lives in minds; a program is a theory the team holds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://simonwillison.net/tags/cognitive-debt/"&gt;Simon Willison&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;"Understand to participate"; interactive explanations as a technique&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://hojberg.xyz/the-programmer-identity-crisis/"&gt;Simon Højberg&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;The reviewer's lived agony; gloss and the first-layer-of-QC problem&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://tim-schipper.nl/en/blog/reviewing-ai-generated-pull-requests"&gt;Tim Schipper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;"An AI PR has no smell"; witnessed, not reviewed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.blog/ai-and-ml/generative-ai/agent-pull-requests-are-everywhere-heres-how-to-review-them/"&gt;GitHub engineering&lt;/a&gt;, May 2026&lt;/td&gt;
&lt;td&gt;The agent-PR volume data; the failure-mode playbook and the failing-test heuristic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mario Zechner / Steve Yegge / HBR&lt;/td&gt;
&lt;td&gt;The cost: merchants of complexity, burnout, intensified work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://distill.pub/2017/aia/"&gt;Carter &amp;amp; Nielsen&lt;/a&gt;, Distill 2017&lt;/td&gt;
&lt;td&gt;The deep root: augment not automate; cognitive transformation over outsourcing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mark Weiser, 1992&lt;/td&gt;
&lt;td&gt;HUDs over copilots&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Read down the table and you can see the conversation's shape: a diagnosis (understanding is the bottleneck), a symptom (reviewers witness instead of review), a cost (debt and burnout), a reframe (understand to participate), a toolkit (explainers, quizzes, micro-worlds), and a design imperative (build HUDs). My take is simple, and it's the through-line of how I work: lead with curiosity, demand artifacts you can poke, and use these techniques to keep understanding alive when the agents are typing fast. I &lt;a href="../../21/curiosity-at-the-wheel/"&gt;wrote about the curiosity piece recently&lt;/a&gt;, and it's the same instinct here.&lt;/p&gt;
&lt;h2 id="check-your-understanding"&gt;Check your understanding&lt;/h2&gt;&lt;p&gt;Here's the quiz. Five cards, instant feedback, and a live score. This is the speed-regulator technique, turned back on the post itself.&lt;/p&gt;
&lt;iframe src="check-understanding.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;If you missed any, that's the spot where your mental model and the field's are still diverging, and it's exactly where cognitive debt would quietly accrue. Ask, find the gap, close it, and ask again. That loop is the whole point.&lt;/p&gt;
</content></entry><entry><title>Going Bayesian automates your manual data analysis</title><link href="https://ericmjl.github.io/blog/2026/7/23/going-bayesian-automates-data-analysis/" rel="alternate"/><updated>2026-07-23T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:b5435342-e4ea-3224-a128-c7bdd3c50931</id><content type="html">&lt;p&gt;If you've ever spent an afternoon clicking through dose-response curves, flagging outliers by hand, there's a way to make the model do that automatically.&lt;/p&gt;
&lt;p&gt;People usually talk about Bayesian modeling in terms of richer uncertainty estimates, prior knowledge, and posterior distributions. All true, and all worth the switch on their own. But there's a benefit almost nobody talks about. Going Bayesian automates away the manual data analysis you'd otherwise do by hand. And once you experience it, you wonder how you ever worked any other way.&lt;/p&gt;
&lt;p&gt;Let me show you what I mean.&lt;/p&gt;
&lt;h2 id="the-manual-outlier-dance"&gt;The manual outlier dance&lt;/h2&gt;&lt;p&gt;Picture a cell viability assay lab. Each 96-well plate holds five dose-response curves. An eighteen-plate experiment generates ninety curves in a single run, and your group is already eyeing robotics and 384-well formats to multiply throughput further.&lt;/p&gt;
&lt;p&gt;I've seen workflows where a single study generates anywhere from 125 to 375 curves. Each one gets inspected by hand.&lt;/p&gt;
&lt;p&gt;For every curve, there's usually a point or two that looks off. The standard workflow goes like this: copy raw data from the instrument output, paste it into Excel to line up sample names with plate positions, then copy again into GraphPad Prism to plot and fit the curve. Spot a deviating point, hit a shortcut key to mark it as an outlier, exclude it, refit. Then look again. Maybe another point looks suspicious now. Repeat. And pray you didn't misalign a paste somewhere along the way.&lt;/p&gt;
&lt;p&gt;For one curve, it'll be seconds, but multiply this over dozens of plates and you have an unbounded amount of time spent clicking on a UI (if one's lucky) or switching between two applications, Excel and GraphPad Prism (if one is unlucky).&lt;/p&gt;
&lt;iframe src="outlier-dance.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Are you really going to hire more research assistants to click through curves? That approach collapses under its own weight.&lt;/p&gt;
&lt;p&gt;Now, a quick but important distinction. There are two completely different reasons to exclude a data point.&lt;/p&gt;
&lt;p&gt;The first is a documented physical cause. You saw a bubble in the well. The pipette tip fell off. The reagent looked cloudy. These are data integrity issues, and you can justify removing them before you look at any numbers; in fact, you &lt;em&gt;should&lt;/em&gt; document these in your lab records.&lt;/p&gt;
&lt;p&gt;The second reason is that &lt;strong&gt;a point is statistically inconvenient&lt;/strong&gt;. It sits too far from the curve and inflates your error. That second reason is the one I want to talk about, because it's both common and statistically broken.&lt;/p&gt;
&lt;h2 id="two-ways-exclusion-corrupts-your-results"&gt;Two ways exclusion corrupts your results&lt;/h2&gt;&lt;p&gt;Excluding a point because it's statistically inconvenient has two deep problems.&lt;/p&gt;
&lt;p&gt;The first is reproducibility. What's your rule for calling something an outlier? Two standard deviations from the fit? Three? The classic tools, like Grubbs' test or the ROUT method that Prism builds in, hand you a threshold. But where did that threshold come from? It's a cutoff someone picked, not a principle someone derived. Two analysts staring at the same curve flag different points, because the decision rests on an arbitrary line in the sand. You're eyeballing with extra steps.&lt;/p&gt;
&lt;p&gt;Here is curve 11 from the demo above, reviewed by two different analysts. Dilution 4 reads 100%, which looks like it belongs to the upper plateau. But the transition zone expects something closer to 73% at that dilution. Analyst A keeps it; Analyst B removes it. Toggle between them and watch the ID50 shift.&lt;/p&gt;
&lt;iframe src="two-analysts.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Same data, same software, different ID50. The point at dilution 4 sits right at the inflection of the curve, so keeping or removing it swings the estimate of where the curve crosses 50%. Which analyst is right? Both think they are, and without a principled criterion, neither can prove the other wrong.&lt;/p&gt;
&lt;p&gt;The second problem is subtler, and it's the one that actually corrupts your results. When you run a test like Grubbs', the test assumes your data come from a clean normal distribution with at most one contaminating point. But dose-response data violate this assumption doubly. The raw points along a dose-response curve don't come from a single normal distribution at all; they follow a sigmoidal shape. You'd have to apply Grubbs' to residuals after fitting, but those residuals are often heteroscedastic (more variance at certain concentrations than others), which breaks the test's assumptions even further. With only five to eight dilutions per curve, the test has almost no power to distinguish a genuine outlier from noise. So you find the outlier, remove it, then refit your curve as if that contamination never existed. Your uncertainty estimates now lie to you. They don't account for the decision you already made. You've baked the removal into your results without propagating the uncertainty that choice introduced.&lt;/p&gt;
&lt;p&gt;This isn't just my opinion. Karch demonstrated in a paper literally titled "Outliers May Not Be Automatically Removed" that automatic removal invalidates confidence intervals and biases your estimates (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/37104797/"&gt;Karch, 2023&lt;/a&gt;). The APA Publication Manual itself states plainly that omitting troublesome observations to present a more convincing story &lt;strong&gt;is prohibited&lt;/strong&gt; (APA, 2020). And the FDA's &lt;a href="https://www.fda.gov/media/70858/download"&gt;Bioanalytical Method Validation guidance&lt;/a&gt; is even more direct: repeat analysis is acceptable only for &lt;strong&gt;assignable causes&lt;/strong&gt; such as equipment failure, sample processing errors, or poor chromatography. "It looked like an outlier" is conspicuously absent from that list.&lt;/p&gt;
&lt;h2 id="down-weight-don-t-delete"&gt;Down-weight, don't delete&lt;/h2&gt;&lt;p&gt;Here's where Bayesian modeling changes the game. Instead of fitting a curve and deciding afterward which points to throw away, &lt;strong&gt;you choose a likelihood&lt;/strong&gt; (the part of your model that scores how well the data fits) &lt;strong&gt;that handles outliers from the start&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The normal distribution, the bell curve you know, has thin tails. Probability drops off fast as you move away from the center. A single point sitting far from the curve becomes astronomically unlikely under a normal likelihood, and the only way the model can reduce that penalty is to drag the entire fit toward the outlier. One bad point yanks your curve.&lt;/p&gt;
&lt;p&gt;The Student-t distribution fixes this. It looks like a normal bell curve, but with heavier tails. More probability spills into the extremes. So when a point sits way off the curve, the likelihood essentially shrugs and says, "that's rare, but not impossible," and the fit barely moves. The further the point drifts from the trend, the less pull it exerts. Its influence fades gradually rather than being a binary in-or-out decision.&lt;/p&gt;
&lt;p&gt;Here is the elegant part. The Student-t has a parameter called the degrees of freedom (nu) that controls how heavy the tails are. At nu=1, you get the canonical Student-t distribution with the heaviest tails; outliers barely register. As nu increases, the tails thin out, and the distribution converges toward the Normal. At nu=30 or above, they are nearly indistinguishable. You can put a prior (your belief about the parameter before seeing data) on nu and let the data tell you how much outlier tolerance each dataset needs. The model learns its own robustness.&lt;/p&gt;
&lt;p&gt;And critically, the point is never deleted. It still contributes to the fit, just with proportionally less weight. Your data stays intact.&lt;/p&gt;
&lt;p&gt;Enough theory. Let me show you what this actually looks like on a dose-response curve with a real outlier.&lt;/p&gt;
&lt;iframe src="bayesian_curves_demo.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;The blue curve is the Normal likelihood fit; the orange dashed curve is the Student-t. Dotted vertical lines mark each model's ID50 estimate. Shaded bands are 94% credible intervals from 4,000 posterior draws (PyMC/NUTS, 4 chains). The buttons go from nu=1 (the Student-t at its heaviest-tailed, most forgiving of outliers) to nu=30 (nearly indistinguishable from the Normal, least forgiving).&lt;/p&gt;
&lt;p&gt;Here is what to watch for. At nu=1, the outlier is absorbed by the heavy tails; the orange curve and its ID50 stay put, right where they belong. At nu=30, the orange and blue curves nearly overlap; the distribution has thinned into a Normal, and the outlier drags the ID50 leftward just like the blue curve does. The interesting territory is nu=3 to nu=5, where the model starts caring about the outlier but doesn't let it dominate. That is the sweet spot for most assay data.&lt;/p&gt;
&lt;p&gt;But this demo handles one weird point on one curve. What happens when an entire curve is the outlier?&lt;/p&gt;
&lt;h2 id="hierarchy-handles-the-weird-curves"&gt;Hierarchy handles the weird curves&lt;/h2&gt;&lt;p&gt;Heavy tails solve the point-level problem. But some curves are weird all on their own. Maybe one dose-response curve came out noisy across every well, or the shape is off in a way that suggests a systematic issue for that row.&lt;/p&gt;
&lt;p&gt;This is where hierarchical modeling earns its keep. Instead of fitting each curve in isolation, a hierarchical model fits all of them together. Each curve borrows strength from the others through a shared population distribution. A noisy curve with a couple of strange wells gets gently pulled toward the behavior of its neighbors. Statisticians call this partial pooling.&lt;/p&gt;
&lt;p&gt;The philosophy mirrors the heavy tails, just one level up. Heavy tails down-weight weird points. Hierarchy regularizes weird curves. It's the same idea, merely applied at different scales.&lt;/p&gt;
&lt;iframe src="hierarchy-demo.html" width="100%" scrolling="no" onload="this.style.height=this.contentDocument.body.scrollHeight+'px'"&gt;&lt;/iframe&gt;&lt;p&gt;Six curves from the same experiment. Five came out clean; curve 3 had a bad day with noisy wells. Toggle between independent fits (curve 3 flails) and hierarchical fits (curve 3 borrows strength from the population and snaps into shape). The faint dashed line on curve 3 shows where the population says it should sit.&lt;/p&gt;
&lt;h2 id="this-is-what-automation-looks-like"&gt;This is what automation looks like&lt;/h2&gt;&lt;p&gt;Here's where everything converges. Once you write down your hierarchical model with a heavy-tailed likelihood, you fit it to all eighteen plates at once. The sampler runs. Heavy tails handle the outlier points. Hierarchy handles the outlier curves. The whole thing runs without a single human judgment call.&lt;/p&gt;
&lt;p&gt;I built a prototype of exactly this for a cell viability assay workflow.&lt;/p&gt;
&lt;p&gt;Thirty-six plates. Four minutes. On a laptop.&lt;/p&gt;
&lt;p&gt;How would you feel about that instead of spending your afternoon doing it by hand?&lt;/p&gt;
&lt;p&gt;No shortcut keys. No research assistants clicking through Prism. No two-analysts-flag-different-points problem. The model handles the data analysis decisions you used to make by hand, and it makes them in a principled, reproducible, mathematically grounded way.&lt;/p&gt;
&lt;p&gt;This is the hidden benefit of going Bayesian. People talk about richer uncertainty, prior knowledge, posterior distributions. All genuine advantages. But the automation of manual data analysis is the one that compounds as your experiments scale. Every new plate, every new study, every new operator, the model just handles it. You get your afternoons back, and your results become reproducible by construction.&lt;/p&gt;
&lt;p&gt;If this resonates, find a statistician or data scientist who works with Bayesian tools. Show them your curve-fitting workflow and ask: "Can we use a hierarchical model with a Student-t likelihood for this?" If they know Bayesian modeling, they'll know exactly what you mean.&lt;/p&gt;
&lt;p&gt;Going Bayesian isn't just about better statistics. It's about building a data analysis pipeline that scales without breaking, handles messy data without manual intervention, and gives you answers you can trust as your experiments grow. The math does the work so your team can focus on the science.&lt;/p&gt;
</content></entry><entry><title>Curiosity at the wheel</title><link href="https://ericmjl.github.io/blog/2026/7/21/curiosity-at-the-wheel/" rel="alternate"/><updated>2026-07-21T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:7d1dae65-cc42-3158-9631-149d4239a7d4</id><content type="html">&lt;p&gt;I took part in Round 2 of the &lt;a href="https://marimo.io/pages/events/notebook-competition-2"&gt;alphaXiv x marimo competition&lt;/a&gt;. The task was to pick a recent ML paper, build a marimo notebook that teaches it, and submit a five-minute walkthrough. I had two papers I wanted to tackle, BitNet's 1.58-bit weights and JiT's "predict the clean image, not the noise," and a tool I've been enjoying: pair-programming marimo notebooks with a coding agent.&lt;/p&gt;
&lt;p&gt;The real surprise was how much &lt;em&gt;of the paper&lt;/em&gt; I learned by the end of the sprint. I came out the other side understanding both papers more viscerally than if I'd read them quietly for a week. The mechanism was the artifact we were building together. The notebook gave me something to poke at, and poking taught me what reading alone couldn't.&lt;/p&gt;
&lt;h2 id="judgment-is-the-slow-part"&gt;Judgment is the slow part&lt;/h2&gt;&lt;p&gt;Judgment takes years. It's the feel for when a result is wrong, the instinct for which experiment to run next, and the sense that a curve looks fishy before you can articulate why. Those instincts compound slowly, and no agent hands them to you. I &lt;a href="../../../6/17/agents-amplify-expertise-and-ignorance/"&gt;wrote about this recently&lt;/a&gt;, and I still believe it.&lt;/p&gt;
&lt;p&gt;But there's a prerequisite elementary step: onboarding onto the basic vocabulary and mental models of a new paper or field. Things like wrestling with what "straight-through estimator" actually means in your hands, or making sense of why predicting the clean image might be &lt;em&gt;easier&lt;/em&gt; than predicting the noise, when the established literature screams the opposite. That part has always been slow for me, because it's a wrestling match with jargon. I read a sentence, hit a term, go look it up, lose the thread, reread, repeat.&lt;/p&gt;
&lt;p&gt;The onboarding can go faster, even if the judgment can't.&lt;/p&gt;
&lt;h2 id="aesthetics-do-pedagogical-work"&gt;Aesthetics do pedagogical work&lt;/h2&gt;&lt;p&gt;Before I started building, I had the agent study the three winning notebooks from round 1. Partly this was competitive intelligence, to obtain a clear bar for what good looked like. But there was a separate surprise: I genuinely enjoyed reading them! The best ones opened with an HTML hero banner carrying the paper title and a one-line thesis. They paired that with interactive widgets where you dragged a slider and watched a distribution react, which made the concepts tactile. Some of them were also game-like. You'd try to find a hidden structure, and being wrong was part of the play rather than a verdict. And they closed with a summary callout that re-stated the conclusions cleanly.&lt;/p&gt;
&lt;p&gt;What makes these notebooks teach is the loop they put you in: predict, then see whether you were right. The interactive widgets drive that loop. Everything else exists to support it: a hero banner gives you a mental anchor before the math lands, and game-like elements turn being wrong into play, which is where the learning actually lives. When being wrong costs nothing, you try more things, and trying more things is what teaches you.&lt;/p&gt;
&lt;h2 id="pair-programming-the-paper-not-just-the-code"&gt;Pair-programming the paper, not just the code&lt;/h2&gt;&lt;p&gt;I fired up a &lt;a href="https://marimo.io"&gt;marimo&lt;/a&gt; notebook on a &lt;a href="https://molab.marimo.io"&gt;molab&lt;/a&gt; sandbox with an NVIDIA RTX PRO 6000 Blackwell, and pair-programmed with &lt;a href="https://z.ai"&gt;GLM-5.2&lt;/a&gt; through &lt;a href="https://opencode.ai"&gt;opencode&lt;/a&gt;. I had the paper and the questions, but I didn't have the patience or tenacity to push through the initial energy barrier on my own. The agent had both, in infinite measure.&lt;/p&gt;
&lt;p&gt;The shape of the work looked like this. I'd ask a question, sometimes framed as "explain this to me like I'm five." The agent would answer in plain words, and then build something that let me &lt;em&gt;feel&lt;/em&gt; the answer. For BitNet, that meant a weight heatmap crushing 32-bit floats into three colors, a slider that let me drag the precision and watch accuracy react, and a live training race on the GPU between binary, ternary, and full-precision models. For JiT, it meant a manifold game where I dragged noisy points and watched what "predict clean" versus "predict noise" actually did to the targets.&lt;/p&gt;
&lt;p&gt;Here's the thing about ELI5 as an operating mode. When I asked for it, I stopped wrestling with jargon and started wrestling with concepts, analogies, and ideas. That's the match I actually wanted. The jargon-wrestling was always fake difficulty, syntax I had to look up before I could even formulate the real question. ELI5 crushed the syntax barrier and put me in front of the idea immediately.&lt;/p&gt;
&lt;h2 id="the-artifacts-taught-me-what-the-abstract-couldn-t"&gt;The artifacts taught me what the abstract couldn't&lt;/h2&gt;&lt;p&gt;The clearest example came halfway through the BitNet notebook. My mental model of the paper went something like: take a pretrained language model, crush its weights to {-1, 0, +1}, and it still works. That's the one-line version, and it's what you'd believe if you stopped at the abstract.&lt;/p&gt;
&lt;p&gt;I told the agent to demonstrate it on a real pretrained GPT-2. The perplexity went from 51 to 33,000. The output collapsed to "present present present present."&lt;/p&gt;
&lt;p&gt;But that's not what the paper claims! BitNet models are &lt;em&gt;born ternary&lt;/em&gt;. They're trained from scratch with the straight-through estimator, which lets gradients flow through the rounding step during training. Retrofitting ternary onto a model that spent its whole pretraining life tuning exact magnitudes breaks the calibration of every layer at once. The agent showed me, by running the experiment I asked for and getting a result that contradicted my hunch.&lt;/p&gt;
&lt;p&gt;So I asked the obvious follow-up: "If we can't retrofit, what's the fix?" That curiosity became the notebook's novel extension. Fine-tune the retrofitted model with the same straight-through estimator, on a small AI-domain corpus, and see what happens. Across three model families (distilgpt2, Qwen2.5-0.5B, Pythia-410m), perplexity recovered by about two orders of magnitude. The retrofitted text went from repetition to real on-topic English about artificial intelligence. The mechanism that makes BitNet trainable from scratch is also the mechanism that partially rescues a retrofitted model. That extension exists because I was wrong about the abstract first. The agent's experiment surfaced the contradiction; reading alone, I could have stayed wrong indefinitely.&lt;/p&gt;
&lt;p&gt;The second paper, JiT, had its own version of this. The paper's title is "Back to Basics: Let Denoising Generative Models Denoise," and I couldn't find a quantitative denoising benchmark inside it. So I got curious. We trained a small x₀-predictor on CIFAR-10, then used it as a one-pass denoiser on held-out test images. It improved the peak signal-to-noise ratio by up to 22 dB, beating a Gaussian-blur baseline at every noise level. The conceptual payoff was even cleaner: denoising and generation turn out to be the same operation, both projection onto the data manifold. That extension lives in the notebook now, and it exists because I asked "what does the title actually mean in practice?" and the agent helped me build the answer. (This JiT notebook is the one that ended up winning Round 2.)&lt;/p&gt;
&lt;h2 id="five-minutes-forced-clarity"&gt;Five minutes forced clarity&lt;/h2&gt;&lt;p&gt;The competition required a five-minute walkthrough video, and that constraint turned out to be its own learning tool. Five minutes forces you to decide what actually matters. Combine that with ELI5 and you get a forcing function: explain the paper simply, in spoken English, in five minutes, to a viewer who hasn't read it. If you can't, you don't understand it yet.&lt;/p&gt;
&lt;p&gt;The agent drafted a script structured as a guided tour through the cells. I sent it to a panel of other subagent reviewers and they tore it apart. Repetitive, shallow, vague.&lt;/p&gt;
&lt;p&gt;That review was a gift. Rewriting the script with the agent's help forced me to decide, in plain spoken English, what each section actually taught and why it mattered. Every time I hit a sentence I couldn't say cleanly, that was a gap in my own understanding. Vague claims like "it works well" had to survive a real question like "compared to what, at what scale?" Oversells had to be walked back to what the data actually supported. Limitations I'd been quietly hoping no one would ask about had to be named out loud, because a judge (or at least, my thesis committee) absolutely would.&lt;/p&gt;
&lt;p&gt;Using the agent to help me communicate clearly made the script much better, and making the script better made my understanding better. Writing in my own voice, with the cells in front of me, is where the last misunderstandings surfaced. Rehearsing for the presentation and trying to anticipate questions about each cell forced me to be intellectually honest about my own gaps. Reading alone makes it easy to hide those gaps from yourself.&lt;/p&gt;
&lt;h2 id="i-stayed-in-the-driver-s-seat"&gt;I stayed in the driver's seat&lt;/h2&gt;&lt;p&gt;This is the part I want to be careful about, because it's easy to misread. The agent did a lot of the typing. It wrote the training loops, built the custom anywidgets, ran the parameter sweeps, refactored the underscore-prefixed names I hate. If you'd watched the session from the outside, you'd have seen the agent producing most of the code.&lt;/p&gt;
&lt;p&gt;But &lt;em&gt;I was driving&lt;/em&gt;. Every prompt I sent encoded a decision. The agent could type, but it couldn't make the judgment calls. Recognizing that the collapsed retrofit was a problem worth solving was mine, and so was seeing that JiT's blurry generations were a real flaw to fix, not just how diffusion looks (hence my prompt: "figure out how to solve the image blur problem on a loop until you are sure you have solved it"). And catching that the script oversold the recovery took a feel for what honest claim the data would support, which is the kind of judgment that takes years.&lt;/p&gt;
&lt;p&gt;The agent's speed amplified my judgment, exactly &lt;a href="../../../6/17/agents-amplify-expertise-and-ignorance/"&gt;the way I wrote about before&lt;/a&gt;. It compressed the onboarding loop, from a week of reading and rereading into an evening of asking and building and poking. The years I'd spent building judgment finally had leverage.&lt;/p&gt;
&lt;h2 id="your-curiosity-sets-the-direction"&gt;Your curiosity sets the direction&lt;/h2&gt;&lt;p&gt;Your curiosity has to set the direction. The agent is patient, the agent is fast, the agent will happily build whatever you ask for and a dozen things you didn't. The discipline is yours. Decide what you actually want to understand. Ask for it in plain words. Demand artifacts you can poke, not paragraphs you can read. When the artifact contradicts your hunch, follow the contradiction; that's where the learning lives. And when you think you understand the whole thing, write the five-minute spoken script. The gaps will announce themselves.&lt;/p&gt;
&lt;p&gt;Some things genuinely take time; onboarding doesn't have to be one of them. Keep your hands on the wheel.&lt;/p&gt;
</content></entry><entry><title>Choose your own AI policy</title><link href="https://ericmjl.github.io/blog/2026/7/20/choose-your-own-ai-policy/" rel="alternate"/><updated>2026-07-20T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:9f7db405-c380-31d1-8481-ecb3e5bca13a</id><content type="html">&lt;p&gt;I just got back from SciPy 2026, my eleventh year at the conference. The loudest thread across the week was AI in open source, and what maintainers are supposed to do about it.&lt;/p&gt;
&lt;p&gt;The thread surfaced in a Scientific Python community session, run as a Birds of a Feather (BoF). Stefan van der Walt walked us through the project's work: the specification process (SPECS, the Python-community analog of PEPs and NEPs), four developer summits, shared tools like &lt;a href="https://repo-review.readthedocs.io/"&gt;RepoReview&lt;/a&gt; and &lt;a href="https://github.com/scientific-python/spin"&gt;&lt;code&gt;spin&lt;/code&gt;&lt;/a&gt;, and &lt;a href="https://mystmd.org/"&gt;MyST&lt;/a&gt; as the documentation engine behind the next &lt;a href="https://jupyterbook.org/"&gt;Jupyter Book&lt;/a&gt;. Then the floor opened for Q&amp;amp;A.&lt;/p&gt;
&lt;p&gt;Someone raised a worry I haven't been able to shake. As Claude and its cousins make it cheap to spin up a working tool, maintainers build bespoke utilities for themselves and stop contributing back. The room offered an analogy: if everyone builds their own vehicle, the roads stop working. Someone drives a tank, someone else a bicycle, and the shared infrastructure collapses. The cheap-prototype problem makes it worse: the version a person publishes often only works for them, pushing the next person to build their own.&lt;/p&gt;
&lt;p&gt;That conversation was still on my mind when a pull request hit my review queue that week, giving me a cleaner frame for these questions.&lt;/p&gt;
&lt;h2 id="two-flavors-of-ai-in-open-source"&gt;Two flavors of AI in open source&lt;/h2&gt;&lt;p&gt;Henry Schreiner opened &lt;a href="https://github.com/scientific-python/cookie/pull/821"&gt;PR #821 on scientific-python/cookie&lt;/a&gt; to add an AI page to the cookie project, the Scientific Python community's recommended project template. (Henry started the draft by pointing Claude Opus at the issue and reworking the output by hand, which is exactly the meta-lesson of the page itself.) The page opens with a distinction I believe is useful here:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A developer driving an interactive AI harness with a capable model, reading the output, and taking responsibility for the result.&lt;/li&gt;
&lt;li&gt;Low-cost models running unattended in automated systems that mass-produce pull requests.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The first is a power tool. The second is slop. Most maintainer frustration about AI contributions comes from the second case, and most of the value comes from the first. Conflating them produces bad policy.&lt;/p&gt;
&lt;p&gt;I see this conflation constantly. A contributor opens a pull request with a generic "I improved the code" message, the diff is a wall of token churn, and when you ask a question the answer comes back in the same flat model-voice. That is slop with a human attached. The human is the relay, not the author. Reviewing it takes longer than writing the change myself.&lt;/p&gt;
&lt;p&gt;The first flavor is what I want more of. A contributor who uses an agent to engage with my project, reads the result, and defends every line, is using a tool. I want to encourage them. Encouragement only lands if contributors can tell which flavor you're asking for, which means writing it down before they open a PR.&lt;/p&gt;
&lt;h2 id="write-your-stance-down"&gt;Write your stance down&lt;/h2&gt;&lt;p&gt;The heart of Henry's PR is &lt;code&gt;AI_POLICY.md&lt;/code&gt;, a project-level file that tells contributors what you expect. The PR sketches three tiers you can adapt, and a &lt;a href="https://github.com/melissawm/open-source-ai-contribution-policies"&gt;community-maintained list of real policies&lt;/a&gt; from dozens of projects shows what other maintainers have chosen.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;All in.&lt;/strong&gt; AI-assisted contributions are welcome on the same footing as any other, as long as they meet your quality bar and are disclosed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Moderate.&lt;/strong&gt; AI assistance is fine, but the contributor has to show real human involvement and prior buy-in before opening the PR.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minimal.&lt;/strong&gt; AI-generated PRs are discouraged or restricted. Use this if you have limited review capacity.&lt;/p&gt;
&lt;p&gt;Which one fits depends on the project. The point, however, is that you should pick one. A contributor reading your &lt;code&gt;AI_POLICY.md&lt;/code&gt; should know what to expect before they open a PR, and a maintainer closing a PR should be able to point at the policy rather than re-litigate it every time.&lt;/p&gt;
&lt;p&gt;The cookie PR leaves the policy out of the template by default, letting each project select and modify their own. I think that's right. A project with one overworked maintainer is in a different situation from a project with twenty regular contributors, and the policy should reflect that. The maintainer reserves the right to take whatever action fits the project; nobody is paying for this work, which means no contributor is owed a review.&lt;/p&gt;
&lt;h2 id="engage-with-the-human-not-the-bot"&gt;Engage with the human, not the bot&lt;/h2&gt;&lt;p&gt;The policy also has a recommendation for how maintainers should treat contributors, and my favorite passage is tucked inside a callout:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;If you maintain a project: try to engage with the human. If they are willing to interact (and not just type "address the review" into their harness), treat them like a human, even if you also see the AI working on their behalf. They also may use AI to address a language barrier.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The failure mode of an AI-assisted contribution is the human tagging out, which is the power-tool flavor sliding back into slop. A contributor pastes review comments into their harness, pastes the harness's reply back into the PR, and walks away. The maintainer is now arguing with a bot. That's draining, and it's disrespectful of volunteer time.&lt;/p&gt;
&lt;p&gt;The policy handles this directly with one rule: &lt;strong&gt;keep human review human-to-human.&lt;/strong&gt; If an AI is responding on your behalf, say so, with a disclaimer line at the top of the comment. You're accountable for every change you submit. Don't make a maintainer talk to an AI without knowing it.&lt;/p&gt;
&lt;p&gt;Disclosure norms pay off here too. The PR recommends the Linux kernel's &lt;code&gt;Assisted-by: &amp;lt;harness&amp;gt;:&amp;lt;model&amp;gt;&lt;/code&gt; trailer for commits. A convention I first saw Daniel Chen use is the &lt;code&gt;:robot: _AI text below_ :robot:&lt;/code&gt; line for any AI-generated prose in a PR or comment. Knowing which model was used lets a reviewer run a different model family to cross-check the work (the PR calls this the "rubber-duck" technique). It also just respects the maintainer. Hiding your process when you contribute to open source is a small lie, repeated. It stings because it withholds the context a reviewer needs to do their job.&lt;/p&gt;
&lt;h2 id="context-is-the-bottleneck"&gt;Context is the bottleneck&lt;/h2&gt;&lt;p&gt;Both worries, fragmentation and slop, trace back to the same constraint: the cost of producing code dropped faster than the cost of producing context. When anyone can ship a working prototype in an afternoon, the scarce thing is the understanding of what the code should do, why it should do it, and how it fits with what already exists.&lt;/p&gt;
&lt;p&gt;This shows up in every issue triage. The issue that says "the build is broken" with no logs, no repro, no environment info. The PR that addresses a problem the contributor never confirmed anyone had. The agent-generated patch that is locally correct and globally wrong because nobody told it about the other module that depends on the current behavior. In each case the gap is the same: context that lives in someone's head and nowhere in writing.&lt;/p&gt;
&lt;p&gt;What helps is documentation written for an agent to read. The cookie PR includes a section on &lt;code&gt;AGENTS.md&lt;/code&gt;, the cross-harness standard file that captures how your repository works: preferred command runners, architecture notes, conventions, traps. A good &lt;code&gt;AGENTS.md&lt;/code&gt; makes the AI far more effective, because the AI stops guessing. The same instinct applies to issues: write down what you tried, what you expected, what you saw. The agent working from a thorough issue is a different animal from the agent working from a one-liner. Agents are also good at producing that context. A minimal reproducible example is tedious for a human; for an agent that can spin up the environment, run the code, and iterate until the bug fires, it's a short task. Your &lt;code&gt;AI_POLICY.md&lt;/code&gt; can ask contributors to use their agent to produce a self-contained repro first. The maintainer gets something runnable; the contributor finds out whether the bug is real before anyone else spends time on it. On the PR side, prompts that shape agent behavior belong in &lt;code&gt;AGENTS.md&lt;/code&gt;, which travels with the code and reaches every contributor's agent: ask for the smallest diff that fixes one thing, and for stacked PRs when a change needs more. AGENTS.md itself is documentation written for an agent.&lt;/p&gt;
&lt;p&gt;AGENTS.md captures mechanics. Intent is a different kind of context, and it matters just as much. Thomas Caswell's SciPy keynote traced the story of Matplotlib's design decisions, the kind of narrative that lives in maintainers' heads and leaves with them. Writing that story down as &lt;a href="https://loki.ws/code/2026/01/25/the-arrow-of-intent.html"&gt;design docs&lt;/a&gt; (a high-level design for the project, then low-level designs and EARS requirement specs per feature) gives an agent the intent it needs to stay grounded. The docs will shift as the project evolves. That malleability is fine, because the current version still beats an agent guessing.&lt;/p&gt;
&lt;p&gt;Later in Q&amp;amp;A, someone who works with NASA data raised the same point from their own side: they had started asking whether to write their documentation for a chatbot instead of a human. They had started shipping documentation as markdown alongside the rendered HTML, because the markdown is what the chatbot consumes. The rendered page is for a human; the source is for the agent. The lesson generalizes. If you want agents to do good work on your project, give them something to read.&lt;/p&gt;
&lt;h2 id="be-courageous"&gt;Be courageous&lt;/h2&gt;&lt;p&gt;Documentation is what you write for the agent; your AI policy is what you write for your contributors.&lt;/p&gt;
&lt;p&gt;If you maintain an open source project, pick an AI policy from any of the three tiers. Write it down, commit it, link to it from your contributing guide, and enforce it. The action that fits your project is correct. A solo maintainer with no review capacity is within their rights to say "no unsolicited AI PRs, full stop." A healthy project with twenty contributors and review bandwidth can be all-in, and benefit from it. Both are legitimate. What's undesirable is having no policy, drifting, and burning out.&lt;/p&gt;
&lt;p&gt;That burnout lands on you. The action that fits &lt;em&gt;you&lt;/em&gt; is also legitimate. If engaging with AI-assisted contributors drains you, require prior buy-in, close the PR, or step back from the project for a week. Some projects are solely volunteer. Protecting your capacity is part of the job.&lt;/p&gt;
&lt;p&gt;Protecting your capacity and welcoming contributors go together: hold the line against tag-out bots, and give your time to the humans behind real contributions. When a contributor shows up willing to engage, even when their harness did most of the typing, treat them like a human. Read their PR, ask your questions, and teach them the parts they don't know yet. That is how open source onboards people. Thomas Caswell onboarded me to Matplotlib a decade ago by walking me through seventy pull requests over two months. The first four were painful for him, but he did it anyway, and I'm still here. That people connection is invaluable.&lt;/p&gt;
&lt;p&gt;That human connection is what got more valuable this year. The cost of writing code dropped. The value of a maintainer who knows their project, writes down their stance, and engages with the humans who show up rose, because &lt;a href="../../../6/17/agents-amplify-expertise-and-ignorance/"&gt;agents amplify whoever brings the judgment&lt;/a&gt;. That's the work. Go do it, and don't apologize for the policy you choose.&lt;/p&gt;
</content></entry><entry><title>xarray in biology, take 2</title><link href="https://ericmjl.github.io/blog/2026/7/19/xarray-in-biology-take-2/" rel="alternate"/><updated>2026-07-19T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:26747dc5-c6f3-38cd-812a-f71b951d6f63</id><content type="html">&lt;p&gt;One year ago at SciPy 2025 I wrote &lt;a href="https://ericmjl.github.io/blog/2025/7/15/how-to-use-xarray-for-unified-laboratory-data-storage/"&gt;a post&lt;/a&gt; arguing that laboratory data should live in a single xarray &lt;code&gt;Dataset&lt;/code&gt;, with sample IDs as the shared coordinate system. The spark for that post was &lt;a href="https://cfp.scipy.org/scipy2025/talk/AARA39/"&gt;Ian Hunt-Isaak's SciPy 2025 talk&lt;/a&gt; on xarray across biology. The hunch back then was this: store measurements, features, model outputs, and train/test splits in one labeled n-dimensional container, and the index-matching tax disappears.&lt;/p&gt;
&lt;p&gt;At this year's SciPy 2026, I closed the loop! During the tutorial, I wrote an &lt;a href="https://github.com/ericmjl/scientific-python-skills/blob/add-xarray-linked-data-skill/skills/xarray-linked-data/SKILL.md"&gt;&lt;code&gt;xarray-linked-data&lt;/code&gt; skill&lt;/a&gt; for Ian's &lt;code&gt;scientific-python-skills&lt;/code&gt; repo. That hunch from a year ago has now become an agent skill, with Marimo notebook examples to steer agents in the right direction on novel situations. Developing the skill has clarified where the thesis holds and where it still needs work, and that's what I'd like to document here.&lt;/p&gt;
&lt;h2 id="the-thesis-restated"&gt;The thesis, restated&lt;/h2&gt;&lt;p&gt;A biological data package should be one file that is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;slice-able by molecular entity.&lt;/strong&gt; Select one &lt;code&gt;molecule_id&lt;/code&gt;, and every assay that molecule appears in comes along for the ride.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;indexed straight into other datasets.&lt;/strong&gt; Selecting by concentration auto-constrains time to that concentration's administration window, and vice versa.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;storing Bayesian estimates next to the raw lab data.&lt;/strong&gt; For example, a posterior KD summary lives in the same file as the ELISA OD readings it was derived from.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;a single artifact.&lt;/strong&gt; One &lt;code&gt;.zarr&lt;/code&gt; or &lt;code&gt;.nc&lt;/code&gt; file that you can ship, share, and re-open three years later without a README.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;self-documenting&lt;/strong&gt;. It should have metadata attached to it without needing to reference other sources.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;built on a DataTree.&lt;/strong&gt; Real assays produce data on incompatible grids; DataTree gives each assay its own node while they share &lt;code&gt;molecule_id&lt;/code&gt; through inheritance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The end goal behind all six bullets is this: to equip one data scientist to curate data packages across an entire campaign (or several at once). The agent skill carries the mechanical assembly and the coding agent runs it. Throughout all this, judgment stays with the human who has seen the warts before.&lt;/p&gt;
&lt;h2 id="why-datatree"&gt;Why &lt;code&gt;DataTree&lt;/code&gt;&lt;/h2&gt;&lt;p&gt;It took me a year to realize that a single &lt;code&gt;Dataset&lt;/code&gt; was not enough. In last year's example, I used one dataset, and it worked because the example was a relatively benign one. The microRNA study had expression, ML features, model outputs, and splits all sharing the &lt;code&gt;molecule&lt;/code&gt; axis. Everything was compatible.&lt;/p&gt;
&lt;p&gt;Real biology rarely is. A nanoparticle characterization campaign produces data on fundamentally incompatible grids:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;DLS can give you a 100-point time series per molecule per replicate (hydrodynamic diameter over an hour).&lt;/li&gt;
&lt;li&gt;HPLC gives you a 1500-point chromatogram per molecule (signal across a retention-time axis).&lt;/li&gt;
&lt;li&gt;ELISA gives you an 8 by 3 grid per molecule per target (eight concentrations, three replicates).&lt;/li&gt;
&lt;li&gt;Flow cytometry gives you 5000 raw events per molecule per cell line.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Forcing those into one flat &lt;code&gt;Dataset&lt;/code&gt; means padding every array to a shared shape, or splitting into multiple files and losing the single-artifact property. Both options give up what made the pattern attractive in the first place.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;xr.DataTree&lt;/code&gt; is the fix. Each assay becomes a node in a tree. Each node keeps its own grid, but can also inherit a shared coordinate from the root; all nodes inherit the &lt;code&gt;molecule_id&lt;/code&gt; coordinate from the root, so selection by molecule propagates everywhere. Here is the pattern from the skill's &lt;a href="https://github.com/ericmjl/scientific-python-skills/blob/add-xarray-linked-data-skill/skills/xarray-linked-data/references/cross-experiment-linking.md"&gt;cross-experiment-linking reference&lt;/a&gt;:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;xarray&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;as&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;xr&lt;/span&gt;

&lt;span class="n"&gt;root&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;coords&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;mol_ids&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="n"&gt;dls_ds&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;diameter_nm&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;replicate&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;time_s&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;dls_timeseries&lt;/span&gt;&lt;span class="p"&gt;)},&lt;/span&gt;
    &lt;span class="n"&gt;coords&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;replicate&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;r1&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;r2&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;r3&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;time_s&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linspace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;hplc_ds&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;signal_au&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;retention_min&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;hplc_chromatograms&lt;/span&gt;&lt;span class="p"&gt;)},&lt;/span&gt;
    &lt;span class="n"&gt;coords&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;retention_min&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;linspace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1500&lt;/span&gt;&lt;span class="p"&gt;)},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;elisa_ds&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;od450&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;concentration_nm&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;replicate&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;elisa_data&lt;/span&gt;&lt;span class="p"&gt;)},&lt;/span&gt;
    &lt;span class="n"&gt;coords&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;egfr&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;her2&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;cd20&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;concentration_nm&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;logspace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;replicate&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;r1&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;r2&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;r3&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;tree&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DataTree&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;from_dict&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/characterization/dls&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;dls_ds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/characterization/hplc&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;hplc_ds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/bioassays/elisa&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;elisa_ds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The DLS node keeps its time axis. The HPLC node keeps its retention axis. The ELISA node keeps its concentration and target axes. They share &lt;code&gt;molecule&lt;/code&gt; for free through inheritance. You save the whole thing as one file:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;tree&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;to_zarr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;s3://bucket/campaign_2026q3.zarr&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That's the contract. One file, every assay on its natural grid.&lt;/p&gt;
&lt;h2 id="custom-indexes-slice-by-what-you-actually-mean"&gt;Custom indexes - slice by what you actually mean&lt;/h2&gt;&lt;p&gt;Coordinating by &lt;code&gt;molecule_id&lt;/code&gt; is the easy case. The skill's &lt;a href="https://github.com/ericmjl/scientific-python-skills/blob/add-xarray-linked-data-skill/skills/xarray-linked-data/references/custom-indexes.md"&gt;custom indexes&lt;/a&gt; are where Ian's own &lt;a href="https://ianhuntisaak.com/xarray-linked-indexes/"&gt;linked-indexes work&lt;/a&gt; does the heavy lifting.&lt;/p&gt;
&lt;p&gt;Consider a dose-escalation time course. You administer 1 nM, then 10 nM, then 100 nM, then 1000 nM, each over its own time window. Response is measured continuously across the whole experiment. Every wet-lab scientist who has run one of these knows the question that comes next: what was the response during the 100 nM phase?&lt;/p&gt;
&lt;p&gt;In &lt;code&gt;pandas&lt;/code&gt;, that question is a multi-column filter where you remember which time indices belong to the 100 nM phase, and re-derive the answer every time you ask. With &lt;code&gt;DimensionInterval&lt;/code&gt;, a custom defined interval, you register the concentration-to-window relationship once, and the question becomes a one-liner.&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# Each concentration is administered over its own time window.&lt;/span&gt;
&lt;span class="c1"&gt;# dose_conc (nM):  1      10       100      1000&lt;/span&gt;
&lt;span class="c1"&gt;# dose_intervals:  [0,60) [60,120) [120,180) [180,240)&lt;/span&gt;

&lt;span class="n"&gt;dose_ds_linked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dose_ds&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;drop_indexes&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;time&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;dose_conc&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;set_xindex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;time&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;dose_intervals&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;dose_conc&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;DimensionInterval&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# What was the response during the 100 nM phase?&lt;/span&gt;
&lt;span class="n"&gt;dose_ds_linked&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dose_conc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# -&amp;gt; time auto-constrained to [120, 180]&lt;/span&gt;

&lt;span class="c1"&gt;# What concentrations were active during this time window?&lt;/span&gt;
&lt;span class="n"&gt;dose_ds_linked&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# -&amp;gt; dose_conc auto-constrained to 1 nM and 10 nM&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The relationship is encoded once. Every query after that is a single &lt;code&gt;sel&lt;/code&gt; call, in either direction.&lt;/p&gt;
&lt;p&gt;The skill ships working implementations of &lt;code&gt;PeriodicIndex&lt;/code&gt; (for wrapping coordinates like flow-cytometry angles or circadian time), &lt;code&gt;CoordinateTransform&lt;/code&gt; (pixel to micrometers in microscopy, m/z to mass in mass spec), &lt;code&gt;NDIndex&lt;/code&gt; (time-locking dosing events across experiments), and &lt;code&gt;DimensionInterval&lt;/code&gt; (the phased-regimen pattern above). Five notebooks, each runnable with &lt;code&gt;uv run&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This is the layer that turns xarray from "labeled arrays" into "a coordinate system that understands your experiment."&lt;/p&gt;
&lt;h2 id="bayesian-estimates-next-to-the-lab-data"&gt;Bayesian estimates next to the lab data&lt;/h2&gt;&lt;p&gt;In &lt;a href="https://ericmjl.github.io/blog/2025/2/23/reliable-biological-data-requires-physical-quantities-not-statistical-artifacts/"&gt;February 2025 I argued&lt;/a&gt; that archival biological data should be physical quantities with Bayesian uncertainty, not statistical artifacts like p-values. The data package is where that argument lands.&lt;/p&gt;
&lt;p&gt;The skill's &lt;a href="https://github.com/ericmjl/scientific-python-skills/blob/add-xarray-linked-data-skill/skills/xarray-linked-data/references/cross-experiment-linking.md#real-world"&gt;real-world example&lt;/a&gt; puts it directly into practice. The same DataTree holds the raw ELISA OD readings in &lt;code&gt;/bioassays/raw&lt;/code&gt; and the Bayesian-derived estimates in &lt;code&gt;/bioassays/derived&lt;/code&gt;:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;derived&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Dataset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;# 4PL fit results&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;ec50_um&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;           &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;cell_line&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ec50_estimates&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;ec50_std&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;          &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;cell_line&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;ec50_stds&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="c1"&gt;# Bayesian KD estimates with HDI bounds&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;kd_nm&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;             &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;kd_means&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;kd_nm_hdi_3pct&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;    &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;kd_lower&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;kd_nm_hdi_97pct&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;   &lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;molecule&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;kd_upper&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;coords&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;target&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;egfr&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;her2&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;cd20&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;cell_line&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;hepg2&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;hek293&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;tree&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;xr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DataTree&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;from_dict&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;root&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/bioassays/raw&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;bioassay&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;/bioassays/derived&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;derived&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The whole provenance chain sits in one file. The OD reading that produced a KD estimate is one &lt;code&gt;molecule_id&lt;/code&gt; lookup away. The HDI bounds travel with the point estimate, so any downstream model that consumes the KD has its uncertainty in the same array. Three years from now, a new data scientist opening this file sees the measurement, the model that produced the estimate, and the uncertainty, all in the same object. No README archaeology, no Slack digging, no "who ran this analysis" archaeology.&lt;/p&gt;
&lt;h2 id="the-end-goal-one-curator-many-campaigns"&gt;The end goal - one curator, many campaigns&lt;/h2&gt;&lt;p&gt;This is the part I've been turning over for the past year.&lt;/p&gt;
&lt;p&gt;The framing I had in my head when I started the skill was "ship a data package without a data scientist in the room." A wet-lab scientist plus an agent plus a template, and out pops a valid &lt;code&gt;.zarr&lt;/code&gt;. But after a few days of thinking, I realized that framing was too optimistic. Building a data package involves judgment calls a wet-lab scientist has no reason to have practiced. Which axis is the entity axis? Where do derived estimates live, and which uncertainty columns travel with them? Did the agent silently drop the HDI bounds when it merged two nodes? Are the units consistent across assays? Did the custom index actually propagate selection, or does it look correct on synthetic data and break on the real thing?&lt;/p&gt;
&lt;p&gt;The agent does the mechanical work fast. The agent also does the wrong thing fast. You need a human in the loop who has seen the warts before. In a biotech, that human is the data scientist.&lt;/p&gt;
&lt;p&gt;Thus, we gain leverage! Without the skill, one data scientist handcrafts one data package per campaign. The work is mostly &lt;code&gt;pandas&lt;/code&gt; merging, &lt;code&gt;dtype&lt;/code&gt; wrangling, file-format archaeology, and re-deriving the same coordinate decisions every time. With the skill, the data scientist designs the campaign template once: which assays live at which nodes, which coordinate is the entity axis, where the derived estimates go, what the provenance chain needs. The agent does the per-assay assembly. The data scientist reviews the output, catches the warts, iterates the template, and moves to the next campaign.&lt;/p&gt;
&lt;p&gt;One data scientist can now curate many campaigns in parallel. The attention that used to go into pandas goes into judgment: assay design review, coordinate-system choices, validation, catching batch effects, deciding what uncertainty to keep. The agent scales the curator; the curator keeps the judgment.&lt;/p&gt;
&lt;p&gt;This is also why I decided to PR to Ian's &lt;code&gt;scientific-python-skills&lt;/code&gt; repo and not keep it in my own vault. A skill that lives in my vault helps me. A skill that lives in a public, community-owned home gives every biotech's data scientist the same leverage, and lets the curation patterns converge across the field.&lt;/p&gt;
&lt;p&gt;For now, if you're working on multi-assay biological data, the &lt;a href="https://github.com/ianhi/scientific-python-skills/pull/3"&gt;skill&lt;/a&gt; is MIT-licensed and the notebooks run with &lt;code&gt;uv run&lt;/code&gt;. Pull it, adapt it, break it, and tell me what fell over.&lt;/p&gt;
</content></entry><entry><title>Ollama, vLLM, and SGLang on Modal</title><link href="https://ericmjl.github.io/blog/2026/7/1/ollama-vllm-sglang-on-modal/" rel="alternate"/><updated>2026-07-01T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:4cc333cb-2340-3846-b96c-e8adcdbb4d3a</id><content type="html">&lt;p&gt;I had Qwen3.6 running on Ollama, hosted on Modal, and it worked. The model answered questions, the endpoint stayed up, and I wired it into my tools. I had a specific reason for running it: I am co-teaching a deep research agent tutorial at SciPy 2026, and I wanted every attendee to have a fast LLM endpoint for the hands-on sessions, including those who do not have access to a paid LLM provider. Then I tried to use it for real work, and the cold starts drove me up the wall.&lt;/p&gt;
&lt;p&gt;Every time the container scaled to zero and I sent a fresh request, I waited. A minute, sometimes two, watching the spinner. The model was right there in the volume. The GPU was right there, billed by the second. And I was staring at a loading bar because Ollama had to boot its server, open the weight files, and load 17 GB of GGUF into VRAM before it could generate a single token.&lt;/p&gt;
&lt;p&gt;I wanted to know: would a different inference engine fix this? I had heard that vLLM, paired with Modal's GPU snapshots, could restore a warm model from a memory snapshot in seconds instead of reloading from disk. So I ran the experiment. Same model, same GPU, same 4-bit quantization level, same hourly cost. The only variable was the engine.&lt;/p&gt;
&lt;p&gt;The result surprised me. vLLM won on cold starts. It won on everything.&lt;/p&gt;
&lt;h2 id="the-setup-held-constant"&gt;The setup, held constant&lt;/h2&gt;&lt;p&gt;To make this a fair fight, I held everything constant except the engine.&lt;/p&gt;
&lt;p&gt;The model is Qwen3.6-27B, a 27-billion-parameter dense model with a hybrid architecture (Gated DeltaNet layers mixed with standard attention). Both deployments serve the same model at 4-bit precision:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Ollama&lt;/strong&gt;: &lt;code&gt;qwen3.6:27b&lt;/code&gt; in GGUF format, Q4_K_M quantization, 17 GB on disk&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;vLLM&lt;/strong&gt;: &lt;code&gt;cyankiwi/Qwen3.6-27B-AWQ-INT4&lt;/code&gt; in safetensors, AWQ INT4 quantization, 19 GB on disk&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Both run on a single Nvidia L40S with 48 GB of VRAM. At Modal's pricing, that is \$0.000542 per second, or about \$1.95 per hour, for either deployment. Same GPU, same rate.&lt;/p&gt;
&lt;p&gt;Both expose an OpenAI-compatible &lt;code&gt;/v1/chat/completions&lt;/code&gt; endpoint. I wrote one benchmark script that hits both URLs with the same prompt, the same &lt;code&gt;max_tokens&lt;/code&gt;, the same sampling parameters, and measures time to first token, decode throughput, and total latency. Identical workload, identical hardware, different engine.&lt;/p&gt;
&lt;h2 id="warm-performance-where-vllm-pulls-ahead"&gt;Warm performance, where vLLM pulls ahead&lt;/h2&gt;&lt;p&gt;Let us start with the easy case: the container is already running, the model is in VRAM, and I send a request. This is steady-state throughput, the number that matters for sustained use.&lt;/p&gt;
&lt;p&gt;I ran each engine six times, generating 300 tokens, and recorded every data point. The results were remarkably consistent for vLLM and surprisingly variable for Ollama.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th style="text-align:center"&gt;Ollama (Q4_K_M)&lt;/th&gt;
&lt;th style="text-align:center"&gt;vLLM (AWQ-INT4)&lt;/th&gt;
&lt;th style="text-align:center"&gt;Difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Time to first token&lt;/td&gt;
&lt;td style="text-align:center"&gt;0.85 ± 0.43s&lt;/td&gt;
&lt;td style="text-align:center"&gt;0.32 ± 0.04s&lt;/td&gt;
&lt;td style="text-align:center"&gt;62% lower, 11x less variance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per second&lt;/td&gt;
&lt;td style="text-align:center"&gt;24.3 ± 2.1&lt;/td&gt;
&lt;td style="text-align:center"&gt;39.4 ± 0.0&lt;/td&gt;
&lt;td style="text-align:center"&gt;62% faster, rock-steady&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total for 300 tokens&lt;/td&gt;
&lt;td style="text-align:center"&gt;11.4 ± 0.7s&lt;/td&gt;
&lt;td style="text-align:center"&gt;7.9 ± 0.0s&lt;/td&gt;
&lt;td style="text-align:center"&gt;31% faster&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;figure id="benchmark-fig" style="margin:1.6em 0 0.4em"&gt;&lt;div id="benchmark-chart" style="width:100%;max-width:960px;"&gt;&lt;/div&gt;&lt;figcaption style="font-size:0.9em;text-align:center;opacity:0.75"&gt;Points: individual runs (6 per engine). Bars: median. &lt;a href="./benchmark-data.json"&gt;Raw data (JSON)&lt;/a&gt;.&lt;/figcaption&gt;&lt;/figure&gt;&lt;/p&gt;
&lt;script&gt;
(function(){
var ENGINES=[{"name":"Ollama","color":"#E07A5F","x0":1,"data":{"ttft":[1.634,0.5513,0.4853,0.6736,0.6853,1.0419],"tps":[27.1716,25.6511,23.3585,22.4884,21.8061,25.1284],"total":[11.0556,10.5314,11.4449,12.0573,12.4251,11.2296]}},{"name":"vLLM","color":"#3B82F6","x0":2,"data":{"ttft":[0.3202,0.275,0.3823,0.3333,0.2996,0.2962],"tps":[39.4235,39.4225,39.4618,39.4316,39.4211,39.3659],"total":[7.9299,7.8849,7.9846,7.9414,7.9098,7.917]}}];
var METRICS=[{"key":"ttft","yrange":[0,1.8],"yticks":[0,0.5,1.0,1.5],"title":"Time to first token (s)"},{"key":"tps","yrange":[20,42],"yticks":[20,25,30,35,40],"title":"Throughput (tokens/s)"},{"key":"total","yrange":[0,14],"yticks":[0,4,8,12],"title":"Total generation time (s)"}];
var XREF=["x","x2","x3"], YREF=["y","y2","y3"];
var JITTER=[-0.16,-0.10,-0.04,0.04,0.10,0.16];
function med(a){var b=a.slice().sort(function(x,y){return x-y});var n=b.length;return n%2?(b[(n-1)/2]):(b[n/2-1]+b[n/2])/2;}
function traces(){var t=[];METRICS.forEach(function(m,mi){ENGINES.forEach(function(e){var xs=JITTER.map(function(j){return e.x0+j});var ys=e.data[m.key];t.push({x:xs,y:ys,type:"scatter",mode:"markers",xaxis:XREF[mi],yaxis:YREF[mi],name:e.name,legendgroup:e.name,showlegend:mi===0,marker:{color:e.color,size:9,opacity:0.8,line:{width:0}},hovertemplate:"%{y:.2f}&lt;extra&gt;"+e.name+"&lt;/extra&gt;"});var md=med(ys);t.push({x:[e.x0-0.2,e.x0+0.2],y:[md,md],type:"scatter",mode:"lines",xaxis:XREF[mi],yaxis:YREF[mi],legendgroup:e.name,showlegend:false,line:{color:e.color,width:5},hoverinfo:"skip"});});});return t;}
function layout(){var dark=document.body.classList.contains("dark-mode");var ax=dark?"#c9ced6":"#3a4252";var gr=dark?"#39404d":"#e4e8ee";var lay={font:{family:"system-ui,-apple-system,Segoe UI,Roboto,sans-serif",color:ax,size:12},paper_bgcolor:"rgba(0,0,0,0)",plot_bgcolor:"rgba(0,0,0,0)",margin:{l:50,r:14,t:40,b:54},grid:{rows:1,columns:3,horizontalspacing:0.07,pattern:"independent"},showlegend:true,legend:{orientation:"h",y:-0.12,x:0.5,xanchor:"center",font:{size:11}},title:{text:"Ollama vs vLLM on Modal: Qwen3.6-27B, single L40S, 6 runs each",font:{size:14}}};METRICS.forEach(function(m,i){var xs=i===0?"xaxis":("xaxis"+(i+1));var ys=i===0?"yaxis":("yaxis"+(i+1));lay[xs]={range:[0.5,2.5],tickvals:[1,2],ticktext:["Ollama","vLLM"],color:ax,tickfont:{size:11},gridcolor:"rgba(0,0,0,0)",linecolor:ax,zeroline:false};lay[ys]={range:m.yrange,title:{text:m.title,font:{size:11}},color:ax,tickfont:{size:10},gridcolor:gr,linecolor:ax,zeroline:false};});return lay;}
var div=document.getElementById("benchmark-chart");
function render(){Plotly.newPlot(div,traces(),layout(),{responsive:true,displayModeBar:false});}
function boot(){if(window.Plotly){render();}else{var s=document.createElement("script");s.src="https://cdn.plot.ly/plotly-2.35.2.min.js";s.onload=render;document.head.appendChild(s);}}
if(document.readyState!=="loading"){boot();}else{document.addEventListener("DOMContentLoaded",boot);}
new MutationObserver(function(){if(window.Plotly){render();}}).observe(document.body,{attributes:true,attributeFilter:["class"]});
})();
&lt;/script&gt;&lt;p&gt;vLLM generates 60% more tokens per second than Ollama on the same GPU. That is a large gap for the same model at the same precision on the same hardware. But the more striking difference is the variance. vLLM's throughput is rock-steady: 39.4 tokens per second on every single run, with a standard deviation of essentially zero. Ollama bounces between 22 and 27 tokens per second run to run. Its time to first token swings from half a second to 1.6 seconds. vLLM always answers in about a third of a second.&lt;/p&gt;
&lt;p&gt;For interactive use, where you feel every hiccup, that consistency matters as much as the raw speed.&lt;/p&gt;
&lt;p&gt;The reason comes down to CUDA graphs. vLLM, when allowed to capture CUDA graphs during startup, batches kernel launches and eliminates Python-level overhead between decode steps. Ollama's llama.cpp backend is well-optimized, but it does not use CUDA graph capture. For this hybrid DeltaNet architecture, where every decode step runs a mix of Triton linear-attention kernels and standard attention kernels, the graph capture matters even more. Each step touches many small kernels, and launching them individually adds up.&lt;/p&gt;
&lt;p&gt;I learned this the hard way. My first vLLM deploy used &lt;code&gt;--enforce-eager&lt;/code&gt;, which disables CUDA graphs. vLLM ran at 12 tokens per second. Slower than Ollama. I removed the flag, vLLM captured its graphs during warmup, and throughput jumped to 44. That one flag was the difference between "why did I bother" and "this is clearly better."&lt;/p&gt;
&lt;h2 id="cold-starts-the-whole-reason-i-started"&gt;Cold starts, the whole reason I started&lt;/h2&gt;&lt;p&gt;Now the interesting part. A warm container is easy. The question I actually cared about was: how long do I wait when the container has scaled to zero and I send the first request?&lt;/p&gt;
&lt;p&gt;Ollama has no snapshot mechanism that works reliably on Modal. I tried Modal's memory snapshots with Ollama early on, and the restored container could not find its model. Ollama runs as a Go server that spawns separate runner subprocesses to hold the GPU context, and those subprocesses do not survive a memory snapshot. The restored server wakes up, tries to talk to a runner that no longer exists, and fails. So every Ollama cold start is a full reload: boot the server, open the GGUF files from the volume, stream 17 GB into VRAM, run a warmup generation.&lt;/p&gt;
&lt;p&gt;vLLM is a different story. It runs as a single Python process, and Modal's GPU snapshot can capture its full state: the loaded weights, the compiled CUDA graphs, the KV cache memory pool, even the JIT-compiled Triton kernels. On cold start, Modal restores the snapshot and vLLM picks up where it left off.&lt;/p&gt;
&lt;p&gt;The cold-start numbers, measured as time from request to first token with the container scaled to zero:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th style="text-align:center"&gt;Cold-start time to first token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Ollama (full reload)&lt;/td&gt;
&lt;td style="text-align:center"&gt;78s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vLLM (GPU snapshot restore)&lt;/td&gt;
&lt;td style="text-align:center"&gt;57s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;vLLM cold starts 27% faster. The 57 seconds is dominated by Modal restoring 19 GB of GPU memory state from the snapshot. That is a bulk memory copy, bounded by hardware bandwidth, and it is still faster than Ollama reading 17 GB from a network filesystem and initializing from scratch.&lt;/p&gt;
&lt;p&gt;I dug into whether 57 seconds could be pushed lower. The honest answer: probably not by much, for a model this size on Modal.&lt;/p&gt;
&lt;p&gt;Modal's documentation is explicit about the limitation. Snapshots help you skip work that is not bottlenecked by storage bandwidth, like library initialization and JIT compilation. But they do not speed up moving weight bytes. At Modal's typical volume bandwidth of one to two GB per second, 19 GB of weights alone takes 10 to 15 seconds before any GPU state restore. Published benchmarks from other developers confirm this floor: a 27B model on an A100 with snapshots and sleep mode gets about 70 seconds cold start. Our 57 seconds on an L40S is actually better than that reference point.&lt;/p&gt;
&lt;p&gt;The only way to get sub-10 second cold starts is to keep a container permanently warm with &lt;code&gt;min_containers=1&lt;/code&gt;. That trades cost (you pay for idle GPU time) for latency. For bursty workloads where you want scale-to-zero, the snapshot restore time is the price of admission.&lt;/p&gt;
&lt;p&gt;Getting the snapshot to work took three rounds of debugging, and each round taught me something specific.&lt;/p&gt;
&lt;h2 id="three-bugs-between-me-and-a-working-snapshot"&gt;Three bugs between me and a working snapshot&lt;/h2&gt;&lt;p&gt;The snapshot did not work on the first try. Or the second. Each failure was a distinct problem with a distinct fix, and I think they are worth walking through because they say something general about deploying inference engines on serverless GPU infrastructure.&lt;/p&gt;
&lt;h3 id="out-of-memory-during-profiling"&gt;Out of memory during profiling&lt;/h3&gt;&lt;p&gt;My first vLLM deploy crashed with a CUDA out-of-memory error. The model loaded fine, but when vLLM tried to profile available memory for the KV cache, it ran out. The GPU showed 42 GB in use with only 2 GB free, on a 48 GB card, before a single token of KV cache was allocated.&lt;/p&gt;
&lt;p&gt;Two things were eating memory. First, I had set &lt;code&gt;--max-num-batched-tokens 131072&lt;/code&gt;, which means vLLM's profiling forward pass runs a dummy batch of 131,072 tokens through the model. For a 27-billion-parameter model, the activation tensors for a batch that size are enormous. Second, vLLM was loading the vision encoder (Qwen3.6 is a multimodal model), which I did not need for text-only inference.&lt;/p&gt;
&lt;p&gt;The fix was &lt;code&gt;--language-model-only&lt;/code&gt; to skip the vision encoder, and reducing &lt;code&gt;--max-num-batched-tokens&lt;/code&gt; to 8192. The model's hybrid architecture helps here: only 16 of the 64 layers use full attention with growing KV cache. The other 48 layers are linear-attention layers with constant memory state. So even with a smaller batched-tokens budget, the effective context capacity is large. After the fix, vLLM reported 21.85 GiB of available KV cache, enough for 332,000 tokens.&lt;/p&gt;
&lt;h3 id="enforce-eager-the-throughput-killer"&gt;Enforce-eager, the throughput killer&lt;/h3&gt;&lt;p&gt;The existing vLLM deployment I was adapting from used &lt;code&gt;--enforce-eager&lt;/code&gt;, a flag that disables &lt;code&gt;torch.compile&lt;/code&gt; and CUDA graph capture. It makes startup faster and uses less memory, which is why it is common in quick-start examples. But it wrecks decode throughput.&lt;/p&gt;
&lt;p&gt;With &lt;code&gt;--enforce-eager&lt;/code&gt;: 12 tokens per second. Without it: 44 tokens per second. Nearly four times faster. The CUDA graph capture happens during the snapshot build, so it costs nothing at cold-start time. The snapshot preserves the captured graphs. Every restored container inherits them for free.&lt;/p&gt;
&lt;p&gt;If you are deploying vLLM on Modal with snapshots, remove &lt;code&gt;--enforce-eager&lt;/code&gt;. Let the graphs compile during the snapshot build and ride them forever.&lt;/p&gt;
&lt;h3 id="the-torch-compile-cache-that-broke-the-snapshot-restore"&gt;The torch compile cache that broke the snapshot restore&lt;/h3&gt;&lt;p&gt;The snapshot was building successfully, but the restore kept failing with a cryptic error: &lt;code&gt;vfs.CompleteRestore() failed: failed to complete restore for filesystem type "9p": failed to walk "torch_compile_cache/torch_aot_compile/..."&lt;/code&gt;. Modal's snapshot restore was trying to walk a file on the vLLM cache volume that did not exist in the restored state.&lt;/p&gt;
&lt;p&gt;The problem was a mismatch between the memory snapshot and the volume. The memory snapshot captured the vLLM process with references to torch compile cache files on the &lt;code&gt;vllm-cache&lt;/code&gt; volume. On restore, Modal re-mounted the volume from its latest committed state, and the cache files the process expected were gone. The 9P filesystem walk failed, the restore aborted, and Modal fell back to a full container init. Cold start was 106 seconds instead of 57.&lt;/p&gt;
&lt;p&gt;The fix was to remove the &lt;code&gt;vllm-cache&lt;/code&gt; volume entirely and redirect &lt;code&gt;TORCHINDUCTOR_CACHE_DIR&lt;/code&gt; to &lt;code&gt;/tmp&lt;/code&gt;. The compiled artifacts live in the process memory, captured by the snapshot. The on-disk cache is only for persistence across cold starts, and with snapshots, you do not need that persistence. The snapshot is the persistence. After removing the volume, cold start dropped to 57 seconds.&lt;/p&gt;
&lt;p&gt;I should note one thing I discovered later: the reason the vLLM sleep endpoint returned a 404 in my first attempt was that I had dropped the &lt;code&gt;VLLM_SERVER_DEV_MODE=1&lt;/code&gt; environment variable when adapting the deployment from an existing repo. That variable exposes the &lt;code&gt;/sleep&lt;/code&gt; and &lt;code&gt;/wake_up&lt;/code&gt; HTTP endpoints. With it set, the proper sleep-then-snapshot pattern (offload weights to CPU before snapshotting, then reload on wake) would work, making snapshots more reliable. For this model size, it would not dramatically change the cold-start number, which is bounded by weight-restore bandwidth. But for correctness and smaller models, it matters.&lt;/p&gt;
&lt;h2 id="what-i-actually-configured"&gt;What I actually configured&lt;/h2&gt;&lt;p&gt;The working vLLM deployment is a single file, &lt;code&gt;vllm_endpoint.py&lt;/code&gt;, in my &lt;code&gt;ollama-on-modal&lt;/code&gt; repo. The core of it is straightforward.&lt;/p&gt;
&lt;p&gt;The vLLM serve command:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;vllm&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;serve&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;MODEL_NAME&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--served-model-name&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;qwen3.6-27b&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--host&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;0.0.0.0&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--port&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;VLLM_PORT&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--tensor-parallel-size&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;1&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--enable-sleep-mode&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--max-num-seqs&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;8&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--max-model-len&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;32768&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--max-num-batched-tokens&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;8192&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--gpu-memory-utilization&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;0.90&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--dtype&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;auto&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--reasoning-parser&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;qwen3&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;--language-model-only&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And the Modal class decorator that makes the snapshot work:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nd"&gt;@app&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;vllm_image&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;gpu&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;L40S&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;scaledown_window&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;MINUTES&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;volumes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;/root/.cache/huggingface&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;hf_cache_vol&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;enable_memory_snapshot&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;experimental_options&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;enable_gpu_snapshot&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The &lt;code&gt;enable_memory_snapshot&lt;/code&gt; plus &lt;code&gt;enable_gpu_snapshot&lt;/code&gt; pair is what lets Modal capture and restore the full GPU state. The snap hook starts vLLM, runs three warmup requests to trigger CUDA graph capture and Triton kernel JIT compilation, then Modal snapshots the warm process. On restore, a separate hook confirms the server is responding. No weight reloading, no graph recompilation.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;scaledown_window&lt;/code&gt; is set to 120 seconds. After two minutes of inactivity, the container scales to zero. The next request triggers a snapshot restore, which takes about 57 seconds. For my usage pattern, that is the right tradeoff between cost (no idle GPU billing) and latency (under a minute to first token from cold).&lt;/p&gt;
&lt;h2 id="wiring-it-into-the-tools-i-actually-use"&gt;Wiring it into the tools I actually use&lt;/h2&gt;&lt;p&gt;Once vLLM was deployed and benchmarked, I pointed two things at it.&lt;/p&gt;
&lt;p&gt;My opencode configuration got a new provider entry pointing at the vLLM endpoint's &lt;code&gt;/v1&lt;/code&gt; base URL. I select &lt;code&gt;vllm-qwen36/qwen3.6-27b&lt;/code&gt; and get 44 tokens per second instead of 28.&lt;/p&gt;
&lt;p&gt;The deep research agent tutorial I am building with &lt;a href="https://www.linkedin.com/in/benbatorsky"&gt;Ben Batorsky&lt;/a&gt; for SciPy 2026 also switched over. The tutorial uses llamabot, which wraps litellm, and the config is just three environment variables:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nv"&gt;LLM_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;openai/qwen3.6-27b
&lt;span class="nv"&gt;TUTORIAL_LLM_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://&amp;lt;your-modal-deployment&amp;gt;.modal.run/v1
&lt;span class="nv"&gt;TUTORIAL_LLM_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;vllm-no-auth
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Rather than hand out a live endpoint, the deployments behind these numbers are open source: &lt;a href="https://github.com/ericmjl/ollama-on-modal"&gt;ollama-on-modal&lt;/a&gt;, &lt;a href="https://github.com/ericmjl/vllm-on-modal"&gt;vllm-on-modal&lt;/a&gt;, and &lt;a href="https://github.com/ericmjl/sglang-on-modal"&gt;sglang-on-modal&lt;/a&gt; each stand one inference engine up on Modal, so you can deploy your own.&lt;/p&gt;
&lt;p&gt;One snag: llamabot's &lt;code&gt;StructuredBot&lt;/code&gt; does a client-side capability check and rejects model names it does not recognize, even when the server supports structured output fine. The fix was three lines in the tutorial's &lt;code&gt;llm.py&lt;/code&gt; that call &lt;code&gt;litellm.register_model&lt;/code&gt; to tell litellm the custom model supports &lt;code&gt;response_schema&lt;/code&gt;. After that, both &lt;code&gt;SimpleBot&lt;/code&gt; and &lt;code&gt;StructuredBot&lt;/code&gt; work against the vLLM endpoint.&lt;/p&gt;
&lt;h2 id="what-i-would-do-differently"&gt;What I would do differently&lt;/h2&gt;&lt;p&gt;If I were starting from scratch, I would skip the Ollama-on-Modal step entirely. Ollama is wonderful for local, single-user inference on a laptop. Its one-command &lt;code&gt;ollama run&lt;/code&gt; experience is unmatched for development and quick experiments. But for a serverless GPU deployment where cold starts matter, vLLM's snapshot compatibility is the deciding factor. Ollama's subprocess architecture fights the snapshot mechanism. vLLM's single-process design cooperates with it.&lt;/p&gt;
&lt;p&gt;The one thing Ollama has going for it in this comparison is simplicity. The deployment is fewer lines of code, the GGUF model format is one file, and the API is clean. If cold starts do not matter for your use case (say, you keep the container warm with a ping), Ollama is perfectly fine and easier to operate.&lt;/p&gt;
&lt;p&gt;But if you are paying for GPU by the second and you want the container to scale to zero between requests, the snapshot restore is the feature that makes that practical. And right now, only vLLM plays nice with it.&lt;/p&gt;
&lt;p&gt;I did not test SGLang in this round. I scaffolded a separate &lt;a href="https://github.com/ericmjl/sglang-on-modal"&gt;sglang-on-modal&lt;/a&gt; repo and deployed it, but SGLang's current release hits a dtype compatibility error with Qwen3.6's hybrid DeltaNet architecture. The model loads and CUDA graphs capture successfully, but inference crashes with a BFloat16 versus Float16 mismatch in the linear attention layers. vLLM 0.23.0 handles this architecture correctly. The SGLang repo is ready for when they ship a fix. Future work: once SGLang supports this architecture, benchmark how it handles Qwen3.6-27B against vLLM on the same GPU and snapshot setup.&lt;/p&gt;
&lt;p&gt;For now, vLLM with CUDA graphs and GPU snapshots gives me 39 tokens per second warm and 57 seconds to first token cold, on a \$1.95-per-hour L40S, and that is the deployment I am keeping.&lt;/p&gt;
</content></entry><entry><title>Agents amplify expertise and ignorance</title><link href="https://ericmjl.github.io/blog/2026/6/17/agents-amplify-expertise-and-ignorance/" rel="alternate"/><updated>2026-06-17T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:4638b6d0-6a24-323d-9458-c67547f9ce13</id><content type="html">&lt;p&gt;I gave a talk on 10 June at Data-Driven Pharma (East) 2026 where I built a Bayesian hierarchical model of protein melting points live, in front of an audience, in 27 minutes (including live questions). A coding agent, Cursor, wrote most of the code in a marimo notebook, paired using Marimo pair. I narrated, took questions while the model sampled, and we landed on a posterior over the melting temperature of a yeast protein.&lt;/p&gt;
&lt;p&gt;The setup -- featuring Cursor and Marimo -- was the message. I really hate making slides, so I decided a live demo was the best thing to do. I opened a &lt;code&gt;marimo&lt;/code&gt; notebook on port 2720, pointed Cursor's agent at it, and started talking to my data the way I actually do at work.&lt;/p&gt;
&lt;p&gt;Halfway through, something I already believed got sharper. The agent was fast. Genuinely fast. But it was fast &lt;em&gt;for me&lt;/em&gt; in a way it would have been merely noisy for someone who lacked the context. Every minute I saved, I saved because of judgment I'd built over years. The agent amplified that judgment.&lt;/p&gt;
&lt;p&gt;Here is what convinced me.&lt;/p&gt;
&lt;h2 id="every-prompt-hid-a-decision"&gt;Every prompt hid a decision&lt;/h2&gt;&lt;p&gt;People treat prompting like a discrete skill, a thing you can grind at in isolation. Read the prompts I gave the agent during the demo and you'll see the real substrate. Each one handed over a decision a scientist had to make first.&lt;/p&gt;
&lt;p&gt;To pick a protein, I had to know that a melting curve only means something when the signal has real dynamic range; it has to actually fall as the protein denatures. So I told the agent:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;Find me a protein with a very good dynamic range between the upper and lower bound signal. If the thing doesn't melt over temperature, or is very labile and melts straight away, that's not something I want to start with.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That is a Bayesian's judgment wearing a prompt's clothes. The agent searched the dataframe. I made the call about what counts as a fittable curve.&lt;/p&gt;
&lt;p&gt;To fit the curve, I told it the four-parameter logistic's upper and lower bounds had to land at exactly &lt;code&gt;1.0&lt;/code&gt; and &lt;code&gt;0.0&lt;/code&gt;. That is dataset-specific domain knowledge, available to me only because I know something about how that data was generated. In plenty of assays the plateaus are sloppy and you let the data find them. In this meltome assay the fold change is normalized, so the bounds are fixed. Only someone who has stared at this assay would know that.&lt;/p&gt;
&lt;p&gt;The hierarchical model at the end carried the same fingerprint. I told the agent to put the population prior on the melting temperature and let the slope be whatever. That choice encodes a scientific belief: proteins in the same organism share a distribution of stability, while their transition steepness stays idiosyncratic. You earn that belief by fitting a lot of curves and watching what breaks.&lt;/p&gt;
&lt;p&gt;Strip the expertise out and the prompts collapse. A novice asks the agent to "analyze the protein data" and gets back a plausible, confident, wrong answer.&lt;/p&gt;
&lt;h2 id="i-caught-what-the-agent-missed"&gt;I caught what the agent missed&lt;/h2&gt;&lt;p&gt;The cleanest moment came early. The agent loaded the parquet and showed me the dataframe. I looked at the columns and noticed the temperature was missing, the one column a meltome analysis needs to exist. I'd prepped that file from a JSON that morning and dropped it on the floor.&lt;/p&gt;
&lt;p&gt;The agent was perfectly happy. It had data, it had columns, it would have charged ahead and fit curves to nothing. I had to send it hunting through the original paper's methods section for the actual temperatures.&lt;/p&gt;
&lt;p&gt;That is the whole game. The agent has blind spots. Left alone, it charges right past them. The person who has run the assay feels the missing column immediately. And the faster the agent moves, the more you need exactly that feeling, because the agent accelerates regardless of whether it is right.&lt;/p&gt;
&lt;h2 id="know-your-vibe-zone"&gt;Know your vibe zone&lt;/h2&gt;&lt;p&gt;Here is the rule I keep coming back to. &lt;strong&gt;Go fast inside your zone of expertise; slow way down outside it.&lt;/strong&gt; In other words, know your &lt;strong&gt;vibe zone&lt;/strong&gt;, the place where your judgment is reliable enough that you can let the agent run and trust your gut-check on the output.&lt;/p&gt;
&lt;p&gt;The four-parameter logistic lives in my vibe zone. I have fit a lot of them. I can glance at a &lt;code&gt;PyMC&lt;/code&gt; model, see a thousand tuning samples and &lt;code&gt;target_accept=0.95&lt;/code&gt;, and know it is behaving. I let the agent write it and skimmed the result.&lt;/p&gt;
&lt;p&gt;Analytical chemistry sits on the other side. I lack proper training in it. Hand me an LC-MS dataset with the same agent and I slow to a crawl, one plot at a time, five minutes of thinking before each request, because a sane result and a broken one look the same to me at a glance.&lt;/p&gt;
&lt;p&gt;The agent widens this difference. Inside my vibe zone it turns ten hours into one. Outside it, it turns confusion into confident confusion, faster. The leverage scales with how much you already know.&lt;/p&gt;
&lt;h2 id="trust-scales-with-your-ability-to-verify"&gt;Trust scales with your ability to verify&lt;/h2&gt;&lt;p&gt;Mid-demo I said something that made the room laugh, and I meant it: I'm going to trust what the agent says. You shouldn't.&lt;/p&gt;
&lt;p&gt;I could trust it because I could check it. I know what a posterior over a melting temperature should look like. I know what a forest plot rank-ordered by probability of superiority does when the data is good, and what it does when the curve is junk. Trust is a dial you set with your ability to verify. Experts turn it up. Beginners keep it low and check everything, which is exhausting, which is why beginners move slowly even with a fast agent.&lt;/p&gt;
&lt;h2 id="the-gap-widens"&gt;The gap widens&lt;/h2&gt;&lt;p&gt;The appealing story about coding agents is that they democratize the craft, that now anyone can build a Bayesian model. My spicy take is that the lived reality runs the other way. The agent hands a loaded instrument to whoever is holding it. In expert hands it is a microscope. In novice hands it is a confocal microscope pointed at the wrong slide, producing beautiful, expensive, wrong images.&lt;/p&gt;
&lt;p&gt;If you have spent a decade learning your domain, this is excellent news. Your investment compounds now instead of depreciating. The boring middle of data science, the boilerplate and plotting and fiddly dataframe wrangling, burns away, and what remains is the part that was always worth something: deciding what to measure, what to believe, and what to distrust.&lt;/p&gt;
&lt;p&gt;The agent made my years matter more. It also raises questions on how to bring people up to speed in the same domain. And yet, if you want more leverage from tools like these, the highest-return thing you can do is still, somewhat against the spirit of the age, to become an expert in something real.&lt;/p&gt;
</content></entry><entry><title>My coding agent learned a lesson and patched its own skill</title><link href="https://ericmjl.github.io/blog/2026/6/16/my-coding-agent-learned-a-lesson/" rel="alternate"/><updated>2026-06-16T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:34762006-75aa-330a-8ea1-c8fd238d2d38</id><content type="html">&lt;p&gt;I was debugging a transcript duplication bug in my voice-first gym coaching app. The coach's responses were being saved twice to the database, one turn apart. I traced it to a React 18 batching issue in the flush logic, refactored the state management into a hook, wrote a one-shot backfill to clean up the historical data, and committed everything.&lt;/p&gt;
&lt;p&gt;then I went to get a coffee.&lt;/p&gt;
&lt;p&gt;When I came back, the skill file for &lt;a href="https://www.convex.dev/"&gt;Convex&lt;/a&gt; migrations had a new entry. The &lt;code&gt;convex-migration-helper&lt;/code&gt; skill, installed in my repo a week ago and untouched since, now contained a six-line callout explaining that &lt;code&gt;internalMutation&lt;/code&gt; functions cannot be invoked from the Convex CLI. The code example had been corrected from &lt;code&gt;internalMutation&lt;/code&gt; to &lt;code&gt;mutation&lt;/code&gt;. A reference file deep in the skill's &lt;code&gt;references/&lt;/code&gt; directory had been updated with the same fix.&lt;/p&gt;
&lt;p&gt;Nobody told it to do that. In another session, my coding agent, GLM-5.2 on OpenCode, hit the &lt;code&gt;internalMutation&lt;/code&gt; wall during the backfill and solved the problem on its own; a background review process then extracted the lesson and patched the skill.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/ericmjl/opencode-autolearn"&gt;&lt;code&gt;opencode-autolearn&lt;/code&gt;&lt;/a&gt; is the plugin that made that possible. It is open source and runs on your machine. I shipped the first version a few weeks back, and it has been compounding ever since. As I write this, across twenty-plus projects it has logged over a thousand observations, spawned nearly three thousand review sessions, and grown a store of fifty-plus skills. My coding agent gets better every session, and I do &lt;em&gt;nothing&lt;/em&gt; extra to make that happen.&lt;/p&gt;
&lt;h2 id="agents-that-forget"&gt;Agents that forget&lt;/h2&gt;&lt;p&gt;I wrote about &lt;a href="../../../1/17/how-to-build-self-improving-coding-agents-part-1/"&gt;building self-improving coding agents&lt;/a&gt; back in January. The core observation was simple: AI coding agents repeat the same mistakes across sessions because they have no mechanism to learn from corrections. Every session starts from scratch. You re-state the same preferences. You re-correct the same behaviors. You are babysitting a very fast intern.&lt;/p&gt;
&lt;p&gt;The post identified two levers: &lt;code&gt;AGENTS.md&lt;/code&gt; as repository memory, and skills as reusable playbooks. Both work. I use them every day. But the loop was still manual. I had to notice the pattern, decide what to do with it, and write the correction myself. I was doing the learning, then handing the agent the results.&lt;/p&gt;
&lt;p&gt;The question that would not leave me alone: what if the agent could watch its own conversations and extract the lessons itself?&lt;/p&gt;
&lt;h2 id="the-inspiration-from-hermes"&gt;The inspiration from Hermes&lt;/h2&gt;&lt;p&gt;A colleague, Edward Miracco, told me about a coding agent called &lt;a href="https://hermes-agent.nousresearch.com/"&gt;Hermes&lt;/a&gt;. Hermes had a property I found fascinating: it got better at working with you over time. The model weights were the same. But Hermes maintained a persistent memory of corrections, preferences, and workarounds, and it updated that memory as you worked.&lt;/p&gt;
&lt;p&gt;I wanted that for &lt;a href="https://opencode.ai"&gt;OpenCode&lt;/a&gt;, the coding agent I use daily. OpenCode has a plugin system, a skill discovery mechanism, and a session model that captures full conversation histories. All the raw materials were there. The missing piece was the feedback loop: something that watched the conversation, decided what was worth learning, and wrote the lessons down.&lt;/p&gt;
&lt;p&gt;So I asked OpenCode to build it.&lt;/p&gt;
&lt;h2 id="what-it-does"&gt;What it does&lt;/h2&gt;&lt;p&gt;&lt;code&gt;opencode-autolearn&lt;/code&gt; is an OpenCode plugin that does four things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Monitors conversations.&lt;/strong&gt; A JavaScript plugin hooks into OpenCode's event system. It counts turns, buffers messages (with secret redaction), and watches for idle periods and session exits.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Spawns review agents.&lt;/strong&gt; Every five assistant turns, or when the session goes idle, or when you close the terminal, the plugin spawns a detached subprocess that runs a review agent. The review agent receives the buffered conversation and an instruction sheet.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extracts lessons.&lt;/strong&gt; The review agent reads the conversation looking for corrections ("don't do X"), preferences ("I prefer Y"), workarounds that worked, and recurring patterns. For each one it finds, it takes action.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Writes to durable stores.&lt;/strong&gt; The review agent uses a Python CLI to update three things: a persistent memory store (loaded into every future session), a user profile (communication and workflow preferences), and skills (created or patched based on observed patterns).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The architecture is deliberately split: a thin JavaScript plugin that only counts and buffers, and a Python CLI that does the data management. The plugin never blocks the main session. Reviews run in a detached subprocess that the plugin fires and forgets. If the review fails, the conversation is saved to a fallback file for debugging. The main session never knows the difference.&lt;/p&gt;
&lt;h2 id="design-decisions"&gt;Design decisions&lt;/h2&gt;&lt;p&gt;Four decisions shaped everything else.&lt;/p&gt;
&lt;h3 id="the-trigger-needs-no-human"&gt;The trigger needs no human&lt;/h3&gt;&lt;p&gt;The review fires on its own. By default, every five assistant turns, or when the session goes idle, or when you close the terminal. You never type a command or decide when to review. The loop does not wait for you.&lt;/p&gt;
&lt;p&gt;This is the decision I care about most. Other tools use human-triggered commands like &lt;code&gt;/dream&lt;/code&gt; for their reflection step, and those work, until they do not. You remember to invoke them for a week. Then you get busy, you forget, and the learning stops. The trigger is the first thing to go when you have real work to do.&lt;/p&gt;
&lt;p&gt;Making it automatic costs almost nothing. A handful of background subprocesses you never see. What you buy with that is a system that learns from every session, not just the ones where you remembered to pull the lever.&lt;/p&gt;
&lt;p&gt;The trigger itself is just an OpenCode plugin. It hooks into OpenCode's event system, counts turns, and fires reviews. No fork, no separate daemon, no modified binary. You install the plugin and your existing OpenCode setup gains the feedback loop.&lt;/p&gt;
&lt;h3 id="a-registry-behind-the-markdown"&gt;A registry behind the markdown&lt;/h3&gt;&lt;p&gt;I started with a single markdown file. Memory lived in &lt;code&gt;~/.autolearn/memory.md&lt;/code&gt;. User preferences lived in &lt;code&gt;~/.autolearn/user-profile.md&lt;/code&gt;. Skills lived in &lt;code&gt;~/.autolearn/skills/{name}/SKILL.md&lt;/code&gt;. All plain text, all human-readable, all directly loadable as OpenCode instructions.&lt;/p&gt;
&lt;p&gt;I picked markdown deliberately. I considered &lt;a href="https://www.sqlite.org/"&gt;SQLite&lt;/a&gt;; it would be more queryable. But the agent reads these files as context, and OpenCode loads instruction files as plain markdown. Markdown served double duty: it was both the storage and the context injection. If I wanted to see what my agent had learned, I &lt;code&gt;cat&lt;/code&gt; the file. If I wanted to correct a lesson, I edited it.&lt;/p&gt;
&lt;p&gt;That worked, until the file filled up. A single markdown file has a size ceiling. When memory hit a few thousand characters, the oldest entries got silently evicted to fit, and I lost lessons I wanted to keep. Reinforcement was append-only: see the same correction three times, get three duplicate entries.&lt;/p&gt;
&lt;p&gt;So the store moved behind the markdown. Today the durable store is a JSONL registry, &lt;code&gt;memories.jsonl&lt;/code&gt;, with one observation per line and no size cap. The markdown file, &lt;code&gt;memory.context.md&lt;/code&gt;, is now a &lt;em&gt;view&lt;/em&gt;, regenerated from the registry on every session start and after every review. Storage and context are separate concerns. The registry can hold thousands of entries; the composed view surfaces the most relevant ones within a soft character budget. Reinforcement counters live in the registry itself, so a repeated correction bumps a counter instead of duplicating a line.&lt;/p&gt;
&lt;p&gt;The agent still reads markdown. I still open a file to see what it learned. The difference is that the file I read is now generated from something more durable underneath it.&lt;/p&gt;
&lt;h3 id="reviews-run-as-subprocesses"&gt;Reviews run as subprocesses&lt;/h3&gt;&lt;p&gt;Each review runs as a separate &lt;code&gt;opencode run&lt;/code&gt; invocation in a detached subprocess. The plugin sets &lt;code&gt;AUTOLEARN_REVIEWER=1&lt;/code&gt; in the subprocess environment so the review session does not trigger its own reviews (which would create an infinite loop).&lt;/p&gt;
&lt;p&gt;The subprocess approach has three benefits. First, failures are isolated: a crashed review does not affect the main session. Second, the review has its own context window: it loads the &lt;code&gt;autolearn-reviewer&lt;/code&gt; skill and gets a clean slate to evaluate the conversation. Third, concurrency is naturally limited to one review at a time, because the plugin tracks an &lt;code&gt;in-process&lt;/code&gt; flag.&lt;/p&gt;
&lt;h3 id="skills-are-symlinked-for-auto-discovery"&gt;Skills are symlinked for auto-discovery&lt;/h3&gt;&lt;p&gt;OpenCode discovers skills in &lt;code&gt;~/.agents/skills/&lt;/code&gt;. When the review agent creates a new skill in &lt;code&gt;~/.autolearn/personas/default/skills/&lt;/code&gt;, the CLI symlinks it into &lt;code&gt;~/.agents/skills/&lt;/code&gt; so OpenCode picks it up automatically. No restart needed, no configuration change.&lt;/p&gt;
&lt;p&gt;This means the loop is: review agent observes a pattern, creates a skill, and the very next session can load that skill if the pattern recurs. The feedback loop closes itself.&lt;/p&gt;
&lt;h2 id="the-build-history"&gt;The build history&lt;/h2&gt;&lt;p&gt;The first commit was a working plugin with the full CLI. I had been thinking about the architecture for a few days, and the initial implementation came out in one piece: turn counting, message buffering, review spawning, memory management, skill creation and patching.&lt;/p&gt;
&lt;p&gt;Then came a series of refinements, each driven by a real problem I hit while dogfooding:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exit-triggered reviews.&lt;/strong&gt; The first version only spawned reviews at turn thresholds. I kept losing the last few turns of a session because I would close the terminal before the threshold fired. So I added &lt;code&gt;beforeExit&lt;/code&gt; and signal handlers (&lt;code&gt;SIGINT&lt;/code&gt;, &lt;code&gt;SIGTERM&lt;/code&gt;) to dispatch a final review on shutdown.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;EARS specifications.&lt;/strong&gt; After the initial build, I had my OpenCode agent write LLDs and EARS specifications for the shipped features. This was partly discipline and partly debugging: the EARS specs allowed me to trace each requirement to actual code paths, and the process surfaced edge cases I had missed. I asked the agent to also add &lt;code&gt;@spec&lt;/code&gt; annotations in the plugin source linking each code block to its EARS requirement ID.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Skill symlinks.&lt;/strong&gt; The initial version created skills in &lt;code&gt;~/.autolearn/skills/&lt;/code&gt; but did not symlink them into &lt;code&gt;~/.agents/skills/&lt;/code&gt;. Skills existed but OpenCode could not discover them. The symlink step closed that gap.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Reinforcement tracking.&lt;/strong&gt; Early on, memory entries were append-only. If the agent observed the same correction three times, it would add three entries. I added a &lt;code&gt;strengths.json&lt;/code&gt; file that tracks how many times each observed pattern has been reinforced, and &lt;code&gt;strengthen&lt;/code&gt;/&lt;code&gt;weaken&lt;/code&gt; commands so the review agent can bump the count instead of duplicating the entry.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Full-text search.&lt;/strong&gt; The review agent needs to answer "has this pattern come up before?" I built an FTS5 index over OpenCode's session database so the review agent can search past conversations before concluding "nothing to record." This catches recurring corrections that were never promoted to memory.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The review-runner wrapper.&lt;/strong&gt; Reviews were leaving behind orphaned sessions in OpenCode's session list. I wrote a shell script wrapper that runs the review, captures the session ID from the JSON output, and deletes the session afterward. The plugin calls the wrapper instead of &lt;code&gt;opencode run&lt;/code&gt; directly.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Behavioral escalation.&lt;/strong&gt; Memory and skills capture in-session lessons, but some corrections recur across every project and belong in the repo's &lt;code&gt;AGENTS.md&lt;/code&gt;, where every agent run reads them. I added a second CLI, &lt;code&gt;improve.py&lt;/code&gt;, that records behavioral rules and counts how often each one recurs. When a rule crosses a threshold, &lt;code&gt;improve.py escalate --apply&lt;/code&gt; writes it into the appropriate &lt;code&gt;AGENTS.md&lt;/code&gt;. The review agent calls &lt;code&gt;improve.py observe ...&lt;/code&gt; as part of every review, so cross-project patterns graduate from session memory into durable repo instructions on their own.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cross-machine sync.&lt;/strong&gt; Everything lived on one machine. Switch laptops, lose the learned memory. I built an E2E-encrypted sync layer: a master password derives a key with PBKDF2-SHA256, the key lives in the OS keychain, and the server only ever sees ciphertext. The plugin auto-syncs on session start and after each review. Two backends implement the same REST API, a self-hosted Fastify server and Convex HTTP Actions, so you can run it yourself or use a hosted one.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Multi-persona stores.&lt;/strong&gt; Work lessons and personal lessons should not collide. I split the store into personas, isolated directories with their own memory, skills, and sync keys. The installer migrates the existing flat layout automatically, so upgrading was invisible.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The memory registry.&lt;/strong&gt; The single &lt;code&gt;memory.md&lt;/code&gt; worked until it filled up and started evicting old entries to fit a size cap. I moved the durable store to a JSONL registry and turned the markdown into a composed view, regenerated on session start. This migration is also where I learned a general lesson the hard way: when you move a system to a new data store, audit every reader of the old one. The autolearn reviewer itself was still warning about "silent 3000-char eviction" long after the cap was gone, because its instruction prose was a stale reader of the old behavior. (That lesson is now in the memory store, funnily enough.)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Curator on a schedule.&lt;/strong&gt; Skills accumulate. Fifty narrow skills are harder to navigate than ten well-named ones. The curator consolidates overlapping skills into broader umbrellas, archives stale ones, and escalates high-reinforcement lessons toward &lt;code&gt;AGENTS.md&lt;/code&gt;. I wired it to the opencode scheduler so it runs weekly without me thinking about it.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-moment-i-knew-it-worked"&gt;The moment I knew it worked&lt;/h2&gt;&lt;p&gt;For the first several days, I was not sure it was working. The plugin was spawning reviews. The observations log was filling up. But I had not seen the system do something I did not expect.&lt;/p&gt;
&lt;p&gt;Then, during a session on my gym-coach project, I hit a wall with the Convex CLI. I had written a backfill mutation as an &lt;code&gt;internalMutation&lt;/code&gt;, tried to invoke it with &lt;code&gt;npx convex run&lt;/code&gt;, and discovered that the CLI can only call &lt;code&gt;mutation&lt;/code&gt;, &lt;code&gt;query&lt;/code&gt;, and &lt;code&gt;action&lt;/code&gt; functions. &lt;code&gt;internalMutation&lt;/code&gt; is private to Convex's internal calling mechanism. I had to convert the function to a regular &lt;code&gt;mutation&lt;/code&gt;, run the backfill, then remove the one-shot code.&lt;/p&gt;
&lt;p&gt;I committed the fix and moved on. The session ended. The review agent spawned.&lt;/p&gt;
&lt;p&gt;When I looked at the &lt;code&gt;convex-migration-helper&lt;/code&gt; skill the next day, it had been patched. The review agent had:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Identified the workaround (convert &lt;code&gt;internalMutation&lt;/code&gt; to &lt;code&gt;mutation&lt;/code&gt; for CLI-invoked backfills).&lt;/li&gt;
&lt;li&gt;Found the existing skill that documented migration patterns (&lt;code&gt;convex-migration-helper&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Patched the skill's &lt;code&gt;SKILL.md&lt;/code&gt; with a new entry explaining when to use &lt;code&gt;mutation&lt;/code&gt; vs &lt;code&gt;internalMutation&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Patched the &lt;code&gt;references/migration-patterns.md&lt;/code&gt; file, correcting the code example and adding a callout box.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The scope was right. It patched the specific reference file where the &lt;code&gt;internalMutation&lt;/code&gt; pattern was documented. It went to the exact section that was wrong and fixed it.&lt;/p&gt;
&lt;p&gt;That was the moment I stopped wondering whether the system worked.&lt;/p&gt;
&lt;h2 id="dogfooding-by-the-numbers"&gt;Dogfooding by the numbers&lt;/h2&gt;&lt;p&gt;I have been running &lt;code&gt;opencode-autolearn&lt;/code&gt; on my machine for a few weeks now. Here is what the system has done in that time, without me lifting a finger:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Observations logged&lt;/td&gt;
&lt;td&gt;~1000 (the log keeps only the most recent thousand)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Review sessions spawned&lt;/td&gt;
&lt;td&gt;~3000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Projects covered&lt;/td&gt;
&lt;td&gt;20+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory entries (registry)&lt;/td&gt;
&lt;td&gt;~100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User profile preferences&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills created&lt;/td&gt;
&lt;td&gt;50+&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Those reviews ran across projects including my gym-coach voice app, my Brain42 knowledge tools, my network analysis teaching materials, my blogbot automation, canvas-chat, and several others. The system watched every session, decided what was worth remembering, and wrote it down.&lt;/p&gt;
&lt;p&gt;The reinforcement tracking is where the compounding shows up. The most-reinforced lesson is a SaaS multi-tenant safety rule: verify before configuring any SaaS service. I hit that pattern across multiple projects and sessions, and each time the review agent bumped the counter instead of adding a duplicate entry. The agent treats that rule as high-priority context because the reinforcement count tells it this one matters.&lt;/p&gt;
&lt;h2 id="what-the-agent-has-learned"&gt;What the agent has learned&lt;/h2&gt;&lt;p&gt;The review agent has created over fifty skills from scratch and patched several existing ones in local repos. Here are a few representative examples:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;code&gt;blogbot&lt;/code&gt; skill for generating social media posts from blog content. The review agent created it after watching me manually draft posts, then patched it with a URL verification step after observing me checking URLs by hand.&lt;/li&gt;
&lt;li&gt;An &lt;code&gt;evergreen-note-quality&lt;/code&gt; skill for my &lt;a href="https://obsidian.md"&gt;Obsidian&lt;/a&gt; vault, created after watching me audit note quality across multiple sessions.&lt;/li&gt;
&lt;li&gt;Bug-pattern skills named after the exact pitfall: &lt;code&gt;react-setstate-in-effect&lt;/code&gt; (don't call &lt;code&gt;setState&lt;/code&gt; unconditionally inside &lt;code&gt;useEffect&lt;/code&gt;), &lt;code&gt;optional-chaining-root-guard&lt;/code&gt; (optional chaining hides a null root), &lt;code&gt;unicode-safe-filename-ops&lt;/code&gt; (normalize before saving a user-titled file), &lt;code&gt;ssrf-guard-node&lt;/code&gt; (validate a URL before a server fetches it). Each one came from a real bug I hit, in a real session, that the review agent watched me fix.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The local repo skill patch is the Convex CLI story I described above: the &lt;code&gt;convex-migration-helper&lt;/code&gt; in the gym-coach repo, patched with the &lt;code&gt;internalMutation&lt;/code&gt; lesson.&lt;/p&gt;
&lt;p&gt;The persistent memory store holds around a hundred entries. Each came from a real mistake. Each has prevented the same mistake in subsequent sessions. The user profile has captured how I like to work: I prefer warm, personal blog conclusions. I expect agents to proactively load writing skills when editing prose. I want autonomous execution without confirmation prompts. I demand quantified evidence in architecture analysis. The agent read these from my conversations and wrote them down. Now every session starts with this context loaded.&lt;/p&gt;
&lt;p&gt;The feedback loop for autolearn looks like this:&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph TD
    A[OpenCode session] --&gt;|every 5 turns / idle / exit| B[autolearn.js plugin]
    B --&gt;|spawn detached subprocess| C[autolearn-reviewer agent]
    C --&gt;|reads conversation| D{Learning opportunity?}
    D --&gt;|correction / preference| E[memories.jsonl registry]
    D --&gt;|recurring pattern| F[create or patch skill]
    D --&gt;|nothing worth recording| G[exit quietly]
    E --&gt;|composed into memory.context.md, loaded into| A
    F --&gt;|symlinked into ~/.agents/skills/| A
&lt;/pre&gt;&lt;p&gt;The core loop works: watch, review, learn, persist, discover. The system improves the agent's behavior without touching model weights.&lt;/p&gt;
&lt;h2 id="install-it-yourself"&gt;Install it yourself&lt;/h2&gt;&lt;p&gt;If you use OpenCode and want to try it:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;curl&lt;span class="w"&gt; &lt;/span&gt;-fsSL&lt;span class="w"&gt; &lt;/span&gt;https://raw.githubusercontent.com/ericmjl/opencode-autolearn/main/install.sh&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;bash
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The plugin activates on your next session. You will not notice it running. But after a few sessions, check &lt;code&gt;~/.autolearn/personas/default/memory.context.md&lt;/code&gt;. Your agent has been taking notes.&lt;/p&gt;
&lt;p&gt;The full source is on GitHub: &lt;a href="https://github.com/ericmjl/opencode-autolearn"&gt;ericmjl/opencode-autolearn&lt;/a&gt;. The design docs include eight LLDs, eleven EARS specifications, and a high-level design that marks every feature shipped, partial, or planned. The README has the complete CLI reference, the sync setup, and the configuration options.&lt;/p&gt;
&lt;h2 id="what-comes-next"&gt;What comes next&lt;/h2&gt;&lt;p&gt;The portability problem is solved. Sync, encryption, and multi-persona stores shipped. The curator runs on a weekly schedule. The memory store grew up too: every entry now carries a retention score and a tier (hot, warm, cold), and the composed view ranks entries by relevance against a soft character budget before it regenerates on each session start. Lessons I keep hitting stay hot; one-off corrections I never repeat fade toward evictable.&lt;/p&gt;
&lt;p&gt;What is genuinely left is calibration, and the one detector that has never run. The retention curve needs real-world tuning; the half-life parameters are fresh guesses, and I want to watch which entries fade too fast or linger too long. The recurring-preference detector, a shift detector that notices when a preference is rising (worth recording) versus settling into habit (learned, stop surfacing), is wired but has never taken a real pass.&lt;/p&gt;
&lt;p&gt;I keep thinking about the moment I saw the patched &lt;code&gt;convex-migration-helper&lt;/code&gt; skill. I had not told anyone to fix it. I had not filed an issue or written a TODO. The conversation where I hit the &lt;code&gt;internalMutation&lt;/code&gt; wall was over. I had moved on. But the system was still watching, and it decided that the workaround I found was worth remembering.&lt;/p&gt;
&lt;p&gt;That is the property I wanted. The agent gets better without me steering the improvement. I do the work, the system does the learning, and the next session inherits the result.&lt;/p&gt;
&lt;p&gt;The Hermes agent had it. Now OpenCode does too.&lt;/p&gt;
&lt;p&gt;I hope autolearn brings you the same quiet compounding it has brought me: an agent that remembers your corrections, respects your preferences, and gets a little sharper every time you sit down to work.&lt;/p&gt;
</content></entry><entry><title>Git worktrees for beginners</title><link href="https://ericmjl.github.io/blog/2026/6/15/git-worktrees-for-beginners/" rel="alternate"/><updated>2026-06-15T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:c7608d5e-7744-3c89-9489-061123ebd887</id><content type="html">&lt;p&gt;You are deep in a feature branch. Files are half-edited, tests are mid-run, your terminal is a graveyard of useful state. Then a teammate pings you: "can you review this PR?" Or a bug on &lt;code&gt;main&lt;/code&gt; needs fixing right now.&lt;/p&gt;
&lt;p&gt;Your options feel lousy. You can stash your work and pray you remember the path back. You can commit a half-finished change just to park it. Or you can clone the repo into a second folder and juggle two copies of the history.&lt;/p&gt;
&lt;p&gt;There is a fourth option, and it is the one I now reach for every time: &lt;code&gt;git worktree&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="what-a-worktree-feels-like-as-a-beginner"&gt;What a worktree feels like as a beginner&lt;/h2&gt;&lt;p&gt;Let's strip away the internals for a second. What is a git worktree? A git worktree shows up as a &lt;strong&gt;new folder on disk that contains a full copy of your repo's files&lt;/strong&gt;, checked out on whatever branch you ask for, placed anywhere you want on your computer.&lt;/p&gt;
&lt;p&gt;That is the essential experience of git worktrees. You run one command, you &lt;code&gt;cd&lt;/code&gt; into a new directory, and you have a clean working tree ready to go. When you are done, you remove it, and the folder disappears.&lt;/p&gt;
&lt;p&gt;Here is the basic loop:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# create a new folder ../hotfix with the repo checked out on a new branch &amp;quot;hotfix&amp;quot;&lt;/span&gt;
git&lt;span class="w"&gt; &lt;/span&gt;worktree&lt;span class="w"&gt; &lt;/span&gt;add&lt;span class="w"&gt; &lt;/span&gt;../hotfix

&lt;span class="c1"&gt;# go work there&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;../hotfix
&lt;span class="c1"&gt;# ...edit, commit, push...&lt;/span&gt;

&lt;span class="c1"&gt;# when you are done, clean it up&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-
git&lt;span class="w"&gt; &lt;/span&gt;worktree&lt;span class="w"&gt; &lt;/span&gt;remove&lt;span class="w"&gt; &lt;/span&gt;../hotfix
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;code&gt;git worktree remove&lt;/code&gt; deletes the folder and deregisters it from git. That is cleanup. If you ever do it by hand (&lt;code&gt;rm -rf ../hotfix&lt;/code&gt;), git leaves a stale registration behind, and &lt;code&gt;git worktree prune&lt;/code&gt; sweeps that up.&lt;/p&gt;
&lt;p&gt;You can have several worktrees at once. &lt;code&gt;git worktree list&lt;/code&gt; shows you every one:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;git&lt;span class="w"&gt; &lt;/span&gt;worktree&lt;span class="w"&gt; &lt;/span&gt;list
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Each entry is a folder on disk, a branch, and a commit. Treat them like separate workspaces that happen to share one brain.&lt;/p&gt;
&lt;h2 id="so-how-is-this-different-from-cloning-into-a-second-folder"&gt;So how is this different from cloning into a second folder?&lt;/h2&gt;&lt;p&gt;That's a great question! If I just need a second copy of the repo, why run &lt;code&gt;git clone&lt;/code&gt; into another folder?&lt;/p&gt;
&lt;p&gt;You can, and people do. I have one colleague at work who has the software developer's equivalent of &lt;code&gt;Untitled12.ipynb&lt;/code&gt; for multiple repos. But a clone and a worktree are doing fundamentally different things under the hood, and the difference shows up the moment you start working.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A clone is a second, independent repository.&lt;/strong&gt; &lt;code&gt;git clone&lt;/code&gt; copies the entire &lt;code&gt;.git&lt;/code&gt; database, history and all, into a brand new repository. The two clones know about each other only as remotes. To share work between them you fetch and push, the same as you would with a colleague's machine. Branches you create in one clone are invisible to the other until you sync them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A worktree is a second door into the same repository.&lt;/strong&gt; It adds a working directory, but the history, the commits, the branches, and the tags all live in one shared &lt;code&gt;.git&lt;/code&gt;. Make a commit in any worktree and it is instantly visible in all the others, because they are reading from the same underlying git database. There is nothing to sync.&lt;/p&gt;
&lt;p&gt;Here is how that plays out side by side:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;git clone&lt;/code&gt; into a new folder&lt;/th&gt;
&lt;th&gt;&lt;code&gt;git worktree add&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Copies the full &lt;code&gt;.git&lt;/code&gt; history again&lt;/td&gt;
&lt;td&gt;Adds only the working files; shares the &lt;code&gt;.git&lt;/code&gt; database&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Branches&lt;/td&gt;
&lt;td&gt;Each clone keeps its own copy; must fetch/push to sync&lt;/td&gt;
&lt;td&gt;All branches shared instantly across worktrees&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New branches&lt;/td&gt;
&lt;td&gt;Invisible elsewhere until pushed&lt;/td&gt;
&lt;td&gt;Visible everywhere immediately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same branch twice&lt;/td&gt;
&lt;td&gt;Allowed (independent checkouts)&lt;/td&gt;
&lt;td&gt;Blocked; a branch lives in one worktree at a time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Network&lt;/td&gt;
&lt;td&gt;Needed to sync between copies&lt;/td&gt;
&lt;td&gt;Never needed (storage is shared locally)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;That last row is the one I feel most. With clones, every "let me grab that branch" is a fetch. With worktrees, the branch is already there because it was never anywhere else.&lt;/p&gt;
&lt;h2 id="the-one-rule-that-trips-people-up"&gt;The one rule that trips people up&lt;/h2&gt;&lt;p&gt;Git enforces a single constraint, and it is worth knowing up front: &lt;strong&gt;a branch can be checked out in only one worktree at a time.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If you are on &lt;code&gt;feature-x&lt;/code&gt; in your main folder, git will refuse to also check out &lt;code&gt;feature-x&lt;/code&gt; in a worktree. That refusal protects you from two working directories racing on the same branch's index. If you need a second working tree for similar work, make a new branch for it:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;git&lt;span class="w"&gt; &lt;/span&gt;worktree&lt;span class="w"&gt; &lt;/span&gt;add&lt;span class="w"&gt; &lt;/span&gt;-b&lt;span class="w"&gt; &lt;/span&gt;feature-x-take2&lt;span class="w"&gt; &lt;/span&gt;../take2
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This is the tradeoff worktrees make. Clones give you total independence at the cost of duplication and syncing. Worktrees give you zero duplication and zero syncing at the cost of that one-branch-one-worktree rule. This is basically perfect for the "park my work and go look at something else" use case.&lt;/p&gt;
&lt;h2 id="are-the-worktrees-linked-on-disk"&gt;Are the worktrees linked on disk?&lt;/h2&gt;&lt;p&gt;Yes, and that is the essence of how git worktrees work.&lt;/p&gt;
&lt;p&gt;A normal repository has a &lt;code&gt;.git&lt;/code&gt; directory that holds everything: the object database (every commit, tree, and blob), the refs (branches and tags), the configuration, the logs. One repository, one &lt;code&gt;.git&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;A linked worktree has a &lt;code&gt;.git&lt;/code&gt; &lt;strong&gt;file&lt;/strong&gt; instead of a &lt;code&gt;.git&lt;/code&gt; directory. That file is one line long:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;gitdir: /Users/you/myproject/.git/worktrees/hotfix
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That path points back into the main repository's &lt;code&gt;.git&lt;/code&gt; directory, into a small admin folder that holds only the per-worktree state: the worktree's own &lt;code&gt;HEAD&lt;/code&gt;, its own &lt;code&gt;index&lt;/code&gt;, a few logs. The heavy lifting, the object database and the refs, is shared from the main repo.&lt;/p&gt;
&lt;p&gt;You can see it for yourself:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;cat&lt;span class="w"&gt; &lt;/span&gt;../hotfix/.git
&lt;span class="c1"&gt;# gitdir: /Users/you/myproject/.git/worktrees/hotfix&lt;/span&gt;

ls&lt;span class="w"&gt; &lt;/span&gt;/Users/you/myproject/.git/worktrees/
&lt;span class="c1"&gt;# hotfix&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That is the linkage. Every worktree points into the same &lt;code&gt;.git&lt;/code&gt; directory, which is why a commit made in one shows up immediately in the rest. They are separate folders on disk with separate working files, but they read and write to one shared repository. One brain, many desks.&lt;/p&gt;
&lt;h2 id="when-i-actually-reach-for-worktrees"&gt;When I actually reach for worktrees&lt;/h2&gt;&lt;p&gt;A few patterns where worktrees have earned their keep for me. Each one is a plain-English ask you can hand to a coding agent:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Quick reviews and hotfixes.&lt;/strong&gt; I am on a feature branch and someone needs me on &lt;code&gt;main&lt;/code&gt;. I &lt;code&gt;git worktree add&lt;/code&gt; a throwaway folder, do the review or the fix, push, and remove it. My feature branch never moves. The ask: "Help me review &lt;code&gt;&amp;lt;PR URL&amp;gt;&lt;/code&gt; in a new worktree."&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Running two versions side by side.&lt;/strong&gt; I keep one worktree on the released version and one on &lt;code&gt;main&lt;/code&gt;, so I can reproduce a bug against both without juggling stashes. The ask: "Check out the v1.2 release tag in a worktree so I can compare it against main."&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Long-running branches with context I want to keep.&lt;/strong&gt; Some experiments take days and accumulate useful terminal state. A worktree lets me park them in their own folder and pop back in without disturbing my main flow. The ask: "Create a worktree for this experiment so I can leave its state alone."&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Handing a whole task to a coding agent.&lt;/strong&gt; A worktree is just a folder, which means I can hand a coding agent the full lifecycle in one instruction: "do this work in a worktree, open a PR, and remove the worktree when you're done." The agent gets its own clean workspace, my main folder sits untouched while it churns, and once the PR is open the throwaway folder disappears on its own. The ease of this is the part I want to flag: the whole loop is small enough that delegating it in plain English feels natural. I am writing this very post that way.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The shared storage means I can have five of these open and still pay only for the working files, five times over, instead of five copies of the entire history.&lt;/p&gt;
&lt;h2 id="a-note-on-cleanup"&gt;A note on cleanup&lt;/h2&gt;&lt;p&gt;&lt;code&gt;git worktree remove &amp;lt;path&amp;gt;&lt;/code&gt; is the polite way out. It deletes the folder and tells git to forget about it. If the worktree has uncommitted changes, git will refuse unless you pass &lt;code&gt;--force&lt;/code&gt;, which is a nice safety net.&lt;/p&gt;
&lt;p&gt;If you ever delete a worktree folder by hand, the registration lingers until you run &lt;code&gt;git worktree prune&lt;/code&gt;. I treat &lt;code&gt;prune&lt;/code&gt; as the "tidy up" command and run it now and then after a week of creating and discarding worktrees.&lt;/p&gt;
&lt;h2 id="give-it-a-try"&gt;Give it a try&lt;/h2&gt;&lt;p&gt;The next time you are tempted to stash or clone just to look at another branch, try this instead:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;git&lt;span class="w"&gt; &lt;/span&gt;worktree&lt;span class="w"&gt; &lt;/span&gt;add&lt;span class="w"&gt; &lt;/span&gt;../scratchpad
&lt;span class="nb"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;../scratchpad
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Or alternatively, ask an agent to work in a worktree for you:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;"Set up a worktree on a new branch and make the changes there."&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Work there. Commit there. When you are done, come back and &lt;code&gt;git worktree remove ../scratchpad&lt;/code&gt;. Your main folder stays exactly as you left it, the throwaway folder is gone, and you never duplicated a single commit to get there.&lt;/p&gt;
&lt;p&gt;One brain, many desks. That is all a worktree is!&lt;/p&gt;
</content></entry><entry><title>Lessons building voice-first AI apps</title><link href="https://ericmjl.github.io/blog/2026/6/13/lessons-building-voice-first-ai-apps/" rel="alternate"/><updated>2026-06-13T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:23eac37f-1362-3125-a882-185fa097ccc6</id><content type="html">&lt;p&gt;Voice-first AI means voice is the primary way you interact with the model. You talk, it talks back, it calls tools, things change on screen. In some of my projects, voice is the &lt;em&gt;only&lt;/em&gt; way you interact with the model. There is no text box.&lt;/p&gt;
&lt;p&gt;That distinction changes everything about how you build, and I learned it through three projects. Yarnsmith is a voice-first game where an AI game master narrates a D&amp;amp;D-style adventure with an educational twist, aimed at parents who want to teach their kids good civic values. Gym Coach is a voice-powered workout companion that talks you through exercises and logs your sets. And I added voice to &lt;a href="https://github.com/ericmjl/canvas-chat"&gt;Canvas Chat&lt;/a&gt;, my visual interface for LLM conversations. Each project surprised me, and the lessons converged on a single pattern: voice plus tools on a web interface equals something genuinely powerful.&lt;/p&gt;
&lt;p&gt;I'm still mapping the edges of this pattern. But I've learned enough to share what works, what doesn't, and where the debugging traps are.&lt;/p&gt;
&lt;h2 id="docs-and-reference-examples-are-non-negotiable"&gt;Docs and reference examples are non-negotiable&lt;/h2&gt;&lt;p&gt;It's tempting to vibe-code a voice app. Spin up the real-time API, talk to the model, assume it works. It rarely does.&lt;/p&gt;
&lt;p&gt;Voice AI is so new that the leading coding models often lack current information about the APIs. Real-time audio streaming, function calling over WebSocket, session management, these are all recent enough that models trained a few months ago will hallucinate the details. They invent parameters that don't exist, use outdated SDK methods, or confidently describe behavior that changed three versions ago.&lt;/p&gt;
&lt;p&gt;The pattern that worked for me is this: get the official documentation, build one working reference example, and point the coding agent at both. Yarnsmith was my first working reference. Once I had a single end-to-end example where the voice pipeline actually functioned, that was enough for Gym Coach and Canvas Chat.&lt;/p&gt;
&lt;p&gt;One working reference does magical wonders; the agents stop guessing and start following.&lt;/p&gt;
&lt;h2 id="build-apis-first-then-layer-on-tool-calls"&gt;Build APIs first, then layer on tool calls&lt;/h2&gt;&lt;p&gt;When I first explored voice as an input modality, it was interesting but limited. The breakthrough for me came when I structured the application around tool calls, like I had seen it in the Claude app with web search as a tool. I started with the &lt;a href="https://ai-sdk.dev/"&gt;Vercel AI SDK&lt;/a&gt; for these TypeScript projects, then switched to the &lt;a href="https://github.com/googleapis/js-genai"&gt;Gemini SDK&lt;/a&gt; when I needed voice-specific features. The architectural pattern stayed the same either way.&lt;/p&gt;
&lt;p&gt;Here is the pattern to follow: Expose all the functionality of your app through API endpoints. Then register those endpoints as tools the voice agent can call. The agent speaks, decides to take an action, calls the tool, and the action happens on the web interface.&lt;/p&gt;
&lt;p&gt;An API gives you flexibility that direct integration never will. You can have the agent call endpoints through tool calls during a voice session. You can also simulate the same interaction manually with &lt;code&gt;curl&lt;/code&gt; to debug individual problems. You can write tests against the same endpoints. The API is the contract, and everything else builds on top of it.&lt;/p&gt;
&lt;p&gt;It's the same layered pattern you see everywhere in software:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A Python library gives you programmatic access&lt;/li&gt;
&lt;li&gt;A CLI wraps the library for command-line use&lt;/li&gt;
&lt;li&gt;An agent wraps the CLI for tool-call use&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Same idea here. You build the API first, then layer on tool calls for the voice agent. The API comes first because it is the thing you can test independently. The tool-call wrapper comes second because it is the thing the agent uses at runtime.&lt;/p&gt;
&lt;h2 id="transcripts-are-your-lifeline"&gt;Transcripts are your lifeline&lt;/h2&gt;&lt;p&gt;When you build a text-based agent, everything is visible. The user types something, the agent responds, you can read the full conversation. When you build a voice agent, most of the interaction is invisible. Audio goes in, audio comes out, and if something goes wrong, you have nothing to look at.&lt;/p&gt;
&lt;p&gt;This is where debugging hell starts.&lt;/p&gt;
&lt;p&gt;The fix is simple but essential: always have a transcript. Every spoken word from the user, every response from the agent, every tool call and its result, all of it should appear as text on the web interface in real time.&lt;/p&gt;
&lt;p&gt;Without a transcript, you are flying blind. You will hear the agent say something unexpected and have no idea what chain of tool calls led to that response. With a transcript, you can trace the exact sequence of events.&lt;/p&gt;
&lt;p&gt;Also, bonus tip: get the agent to give you a &lt;code&gt;Copy&lt;/code&gt; button so that you can copy/paste as much contextual information over to the coding agent's context later!&lt;/p&gt;
&lt;h2 id="log-every-tool-call-with-timestamps"&gt;Log every tool call with timestamps&lt;/h2&gt;&lt;p&gt;A transcript of spoken words is the minimum. Voice agents introduce a problem that text agents don't have: latency.&lt;/p&gt;
&lt;p&gt;Audio processing, network round trips, model inference, tool execution. Each step adds delay. When the agent takes three seconds to respond, you need to know where those three seconds went. Was it the speech-to-text? The model thinking? A slow API endpoint?&lt;/p&gt;
&lt;p&gt;Every entry in your transcript needs a timestamp. Not just the spoken words, but every tool call: when it was initiated, when it completed, what arguments were passed, what was returned. The tool call order matters too, because voice agents can trigger cascading sequences of calls that interact in unexpected ways.&lt;/p&gt;
&lt;p&gt;I guarantee you'll hit latency problems with voice agents. When you do, the timing information in your transcript is the only way to diagnose where the delay lives.&lt;/p&gt;
&lt;p&gt;This principle applies to text agents as well. But it is even more critical for voice because the user is waiting in real time. A three-second delay in a text chat is tolerable. In a voice conversation, it feels broken.&lt;/p&gt;
&lt;h2 id="centralize-voice-operations-into-one-component"&gt;Centralize voice operations into one component&lt;/h2&gt;&lt;p&gt;Good logging tells you what went wrong. Clean architecture determines whether you can fix it. Here is a lesson I learned the hard way with Gym Coach.&lt;/p&gt;
&lt;p&gt;The symptom was specific and maddening. The user would say "sounds good," and the coach would call &lt;code&gt;save_plan&lt;/code&gt; to create a workout. Then silence. The coach stopped talking. The transcript showed the tool call completed, but the voice agent never generated its next response. Gemini was waiting for a &lt;code&gt;FunctionResponse&lt;/code&gt; that never came.&lt;/p&gt;
&lt;p&gt;Debugging this was hell because voice handling was scattered across four components. &lt;code&gt;GeminiLiveClient&lt;/code&gt; managed the WebSocket connection. &lt;code&gt;TurnLifecycle&lt;/code&gt; tracked phantom turns. &lt;code&gt;ToolCallGatekeeper&lt;/code&gt; enforced behavioral rules. The session page held its own callback configuration and workout state. Each component had its own assumptions about what should happen next, and tracing where the response got swallowed took hours. The root cause turned out to be the guards themselves. I had built behavioral guards to prevent bad behavior: premature logging, exercise mismatches, duplicate plans. When a guard fired, it blocked the tool call and discarded the response. But Gemini was waiting for that response. No response meant no next turn. No next turn meant silence.&lt;/p&gt;
&lt;p&gt;Even with logging in place, I would fix one guard in one file and then hit a related bug in another because the same logic was duplicated elsewhere. The architecture was fighting me.&lt;/p&gt;
&lt;p&gt;The fix was architectural, and it happened incrementally. After each debugging session, I asked my coding agent to use the &lt;code&gt;improve-codebase-architecture&lt;/code&gt; skill to find the top recommendation for preventing the class of bug I had just fixed. Instead of a full review process, I said "take your top recommendation and implement it." One recommendation per session. Ship it. Move on. I would strongly encourage this practice for any voice-first project. Each round surfaced a concrete structural fix, like extracting scattered coordination logic into one class, or moving guards from the transport layer to the domain layer. One recommendation at a time kept the changes small enough to verify.&lt;/p&gt;
&lt;p&gt;Over multiple iterations, all voice operations collapsed into a single component, the &lt;code&gt;GeminiLiveClient&lt;/code&gt;. Its contract is clean: it takes a config and a set of event callbacks. One callback, &lt;code&gt;onFunctionCall&lt;/code&gt;, returns a string that becomes the &lt;code&gt;FunctionResponse&lt;/code&gt; sent back to Gemini. The &lt;code&gt;TurnLifecycle&lt;/code&gt; module was deleted entirely. Behavioral guards were stripped out. What remained were data-integrity guards only: dedup duplicate tool calls by ID, rate-limit to one per turn, prevent double-saved plans.&lt;/p&gt;
&lt;p&gt;The key change replaced blocking with description. When a guard fires now, it still sends a &lt;code&gt;FunctionResponse&lt;/code&gt;, but the response says "blocked: plan already saved" instead of going silent. Gemini reads that and self-corrects on its next turn. The coach never stops talking.&lt;/p&gt;
&lt;p&gt;The skeleton of the centralized client looks roughly like this:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kd"&gt;interface&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;VoiceClientEvents&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nx"&gt;onAudioData&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;audio&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;Uint8Array&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="ow"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nx"&gt;onTextResponse&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="ow"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nx"&gt;onInputTranscription&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="ow"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="c1"&gt;// Returns a descriptive string that becomes the FunctionResponse.&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="c1"&gt;// Example: &amp;quot;blocked: plan already saved&amp;quot;, &amp;quot;Set logged. Acknowledge briefly.&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nx"&gt;onFunctionCall&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="ow"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nx"&gt;onTurnComplete&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="ow"&gt;void&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;VoiceClient&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;Session&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;seenToolCallIds&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="ow"&gt;new&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;Set&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;pendingResponses&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}[]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;

&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="kr"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;VoiceConfig&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;Partial&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;VoiceClientEvents&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;

&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="k"&gt;async&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="c1"&gt;// Open the realtime session and register a single message handler&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="k"&gt;private&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;handleMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;ServerMessage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;toolCall&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;fc&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;toolCall&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;functionCalls&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;seenToolCallIds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;fc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;seenToolCallIds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;fc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onFunctionCall&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;.(&lt;/span&gt;&lt;span class="nx"&gt;fc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;fc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pendingResponses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;fc.id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;result&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;??&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ok&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onAudioData&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;.(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onTextResponse&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;.(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="c1"&gt;// ALWAYS flush tool responses. Blocking the FunctionResponse&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="c1"&gt;// means Gemini hangs silently, waiting for a reply that never comes.&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;turnComplete&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pendingResponses&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sendToolResponse&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;functionResponses&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pendingResponses&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;seenToolCallIds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onTurnComplete&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;.();&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The entire contract is one class with one event interface. The domain layer talks to the voice agent through a single callback that returns a string. Tool responses always flush. The only guards that remain prevent duplicate data writes, like logging the same set twice or saving the same plan twice. They never block the voice agent from speaking.&lt;/p&gt;
&lt;p&gt;The results were qualitatively dramatic. Undesirable behaviors that had been recurring were eliminated in the next iteration. Multiple bugs I could foresee were never introduced. New bugs showed up, of course, but the trajectory was a series of step-function improvements, each one making the whole system more debuggable.&lt;/p&gt;
&lt;h2 id="voice-controlled-ui-is-magical"&gt;Voice-controlled UI is magical&lt;/h2&gt;&lt;p&gt;I want to end on the part that made all of this worth building.&lt;/p&gt;
&lt;p&gt;There is something viscerally magical about speaking to your computer and watching it respond. You say "log my set" and the UI updates. You say "show me the next exercise" and the screen changes. It is Geordie La Forge on the Enterprise going "Computer, do something" and watching it happen.&lt;/p&gt;
&lt;p&gt;That experience is genuinely joyful for the user. And it relies on a design principle that is easy to miss: the action needs to be visceral and exposed. When the agent takes a tool call, the user should see the result on screen immediately. The visual context should change in response to the voice command.&lt;/p&gt;
&lt;p&gt;This is where good UI and UX design principles earn their keep. The web interface doubles as the visual feedback layer that makes voice interaction feel real. Every tool call should produce a visible change. Every state transition should be immediate. The user needs to feel that their voice caused something to happen.&lt;/p&gt;
&lt;h2 id="what-i-m-still-figuring-out"&gt;What I'm still figuring out&lt;/h2&gt;&lt;p&gt;The pattern I've converged on, voice plus tools on a web interface, feels powerful but immature. The engineering practices I described above are the scaffolding that makes it reliable enough to ship. Documentation-driven development keeps the models honest. API-first architecture keeps the system testable. Transcripts with timestamps keep the debugging tractable. Centralized voice operations keep the codebase maintainable.&lt;/p&gt;
&lt;p&gt;I'm not yet 100% sure where voice-first AI is heading. But the building experience has been the most fun I've had with software in a while. When you speak to a computer and it responds, when you watch the interface change because you asked it to, you feel like you're living in the future. The trick is making sure that future also has good logging.&lt;/p&gt;
</content></entry><entry><title>What is an agent harness?</title><link href="https://ericmjl.github.io/blog/2026/6/12/what-is-an-agent-harness/" rel="alternate"/><updated>2026-06-12T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:c1485c38-eb41-3632-a94b-4c1577fbc309</id><content type="html">&lt;p&gt;My friend Sean Law asked me this question on LinkedIn recently, essentially, "Do you have a definition of the term agent harness?" I looked around and realized that nobody really has one. &lt;a href="https://simonwillison.net/2025/Sep/18/agents/"&gt;Simon Willison&lt;/a&gt; writes extensively about agents and defines the agent itself ("tools in a loop to achieve a goal") but leaves the harness unspecified. &lt;a href="https://twitter.com/karpathy/status/2024987174077432126"&gt;Andrej Karpathy&lt;/a&gt; calls "Claws" a new layer on top of agents and describes what they do without defining what they are. &lt;a href="https://www.anthropic.com/engineering/building-effective-agents"&gt;Anthropic's guide&lt;/a&gt; to building effective agents is really a guide to harness engineering, but they never use the word. The term is doing a lot of quiet work across the ecosystem. Time to make that work explicit; here is my attempt at doing so.&lt;/p&gt;
&lt;h2 id="defining-the-harness"&gt;Defining the harness&lt;/h2&gt;&lt;p&gt;An agent harness is everything that constrains, shapes, and thus defines what an agent can do. It has four components: the tools you give the agent, the environment those tools run in, the hard controls that architecturally limit what the agent can do, and the soft controls that steer its behavior through prompting.&lt;/p&gt;
&lt;p&gt;Together, these four components define the agent's &lt;strong&gt;action space&lt;/strong&gt;. The action space is the set of all things the agent can actually accomplish. Change any one component and the action space changes.&lt;/p&gt;
&lt;h2 id="hard-controls-vs-soft-controls"&gt;Hard controls vs. soft controls&lt;/h2&gt;&lt;p&gt;Not all constraints work the same way.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hard controls&lt;/strong&gt; are technological constraints on an agent's action space. The agent &lt;em&gt;literally cannot do the thing&lt;/em&gt;. If you do not give a coding agent network access, it cannot phone home. If you do not give a voice agent the ability to speak, it can only respond in text. Hard controls are architectural: they are built into what the agent physically or technologically can and cannot do.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Soft controls&lt;/strong&gt; are the prompted nudges. You write instructions telling the agent to prefer certain behaviors, avoid certain patterns, or follow certain procedures. The agent can still deviate. Soft controls steer; they do not bind.&lt;/p&gt;
&lt;p&gt;In practice, those instructions usually live in one of two places. An &lt;code&gt;AGENTS.md&lt;/code&gt; file is always-on context: repo norms and preferences that ride along with every request. An agent skill file is the on-demand kind: a procedural workflow the agent pulls in only when the task at hand matches it. Both are soft controls; the difference is whether they apply all the time or just when triggered.&lt;/p&gt;
&lt;p&gt;The distinction matters because soft controls are easier to set up but easier to bypass. (More on that later in my voice agent example.) Hard controls require more engineering effort but are more reliable.&lt;/p&gt;
&lt;h2 id="three-harness-examples"&gt;Three harness examples&lt;/h2&gt;&lt;p&gt;Making this concrete helps. Here are three harnesses and how they differ.&lt;/p&gt;
&lt;h3 id="coding-agent-harness"&gt;Coding agent harness&lt;/h3&gt;&lt;p&gt;Modern coding agents like Cursor and OpenCode give you a pretty full toolkit: file read and write, bash execution, code execution, and web search. The environment is your local filesystem (usually sandboxed). The action space is broad; these agents can research, write, and run code all in one session.&lt;/p&gt;
&lt;p&gt;The hard controls are architectural: the agent runs in a sandbox, cannot access your keychain or password manager, and has no ability to escalate privileges. The soft controls are your instructions: prefer Python, follow the testing conventions in this repo, ask before running destructive commands.&lt;/p&gt;
&lt;h3 id="data-science-agent-harness"&gt;Data science agent harness&lt;/h3&gt;&lt;p&gt;A data science agent extends the coding agent harness. It has the same tools (bash, code execution, file read/write, web search) but adds marimo notebooks to the environment. The agent works inside a running notebook kernel, editing cells, executing code, and documenting findings inline.&lt;/p&gt;
&lt;p&gt;The hard controls are nearly identical to a regular coding agent: sandboxed execution, no access to production credentials. The soft controls are where the difference shows up. My instructions to the agent include doing code execution, editing, and documentation all inside the notebook rather than in separate files. The agent could work in standalone &lt;code&gt;.py&lt;/code&gt; files (it has the tools), but I steer it toward the notebook because that is where the analysis lives.&lt;/p&gt;
&lt;h3 id="voicepal-interview-harness"&gt;Voicepal interview harness&lt;/h3&gt;&lt;p&gt;Here is an example that makes the soft control boundary tangible. I use an interview agent inside Voicepal. It is prompted to ask me questions and not give me answers. When it asks me something and I reply, "I don't know man, you tell me," it says, "I'd love to, but my role is to ask you questions and not give you answers."&lt;/p&gt;
&lt;p&gt;That sounds like a hard control, but it is not. I can get around it. I say, "How about you tell me one option and ask me what I think about it?" And it says, "Fair," and then gives me the option. The soft control bent without breaking. I respected the spirit of the constraint (it still asked me a question), but I nudged the agent to give me an answer inside that frame.&lt;/p&gt;
&lt;p&gt;That is the nature of soft controls. They are persuasive, not binding. A determined user can work around them while staying within the letter of the prompt. Hard controls do not have this property. If you do not give the agent a tool, no amount of clever prompting will conjure it.&lt;/p&gt;
&lt;h2 id="bringing-it-together"&gt;Bringing it together&lt;/h2&gt;&lt;p&gt;There is a pattern across all three examples that took me a while to appreciate: the harness is defined by what you include &lt;em&gt;and&lt;/em&gt; what you exclude. Every tool you withhold is as much a design decision as every tool you provide.&lt;/p&gt;
&lt;p&gt;Here is a summary of the three harnesses:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Coding agent&lt;/th&gt;
&lt;th&gt;Data science agent&lt;/th&gt;
&lt;th&gt;Voicepal interview agent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tools&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;File read/write, bash, code execution, web search&lt;/td&gt;
&lt;td&gt;File read/write, bash, code execution, web search, marimo kernel&lt;/td&gt;
&lt;td&gt;Conversation (speech in, text out)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Environment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Sandboxed filesystem&lt;/td&gt;
&lt;td&gt;Sandboxed filesystem + notebook kernel&lt;/td&gt;
&lt;td&gt;Voicepal app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hard controls&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Sandboxed, no keychain access, no privilege escalation&lt;/td&gt;
&lt;td&gt;Same as coding agent&lt;/td&gt;
&lt;td&gt;No tool use beyond conversation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Soft controls&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prefer Python, follow testing conventions, ask before destructive commands&lt;/td&gt;
&lt;td&gt;Do all work inside the notebook, document findings inline&lt;/td&gt;
&lt;td&gt;Ask questions, do not give answers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;An agent harness is the combination of tools, environment, hard controls, and soft controls that together define what an agent can and cannot do. The data science agent shows that two harnesses can share nearly identical hard controls and still produce different behavior, because the soft controls steer the agent in a different direction. The Voicepal example shows that soft controls are persuasion, not architecture; a determined user can work around them.&lt;/p&gt;
&lt;p&gt;I hope this definition brings you clarity the next time someone brings up agent harnesses in conversation. And if they have a different definition, I would love to hear it; the term will only get sharper if more people take a crack at defining it.&lt;/p&gt;
</content></entry><entry><title>How to Multitask Better with Agent Harnesses</title><link href="https://ericmjl.github.io/blog/2026/6/8/how-to-multitask-better-with-agent-harnesses/" rel="alternate"/><updated>2026-06-08T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:08c8e4bb-a6ee-30dd-bd75-2e3dbf115b29</id><content type="html">&lt;p&gt;In my &lt;a href="../../4/exploring-agent-harnesses/"&gt;previous post&lt;/a&gt;, I compared three agent harnesses across workspaces, notifications, automations, and open-source status. Once you have picked a harness, the next question is: how do you actually use it?&lt;/p&gt;
&lt;p&gt;The framework I have landed on is three tiers: one foreground task, one or two background tasks, and however many automated tasks you want running in the shadows. Right now, as I write this post, all three tiers are active. Let me walk through what each one looks like.&lt;/p&gt;
&lt;h2 id="the-three-tiers"&gt;The three tiers&lt;/h2&gt;&lt;h3 id="foreground-your-intellectual-work"&gt;Foreground: your intellectual work&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;Keep one foreground task at a time. This is your intellectual work; protect it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Writing this blog post is my foreground task right now. I am actively reviewing the draft the agent ghostwrote for me: iterating with my agent, refining what I want to say, editing prose, and nudging the agent to further match my voice. There is a meta-layer too: I have a ghostwriter skill that teaches the agent to write like me, and as I review its output, I feed corrections back into the skill itself. The writing improves, and the skill improves, in the same loop.&lt;/p&gt;
&lt;p&gt;This is what foreground work looks like across the board: highly interactive, requiring your full attention. You are pair-coding, pair-writing, or pair-creating with the agent, steering it, reading its traces, deciding whether it is going in the right direction. When the foreground task needs you, everything else waits. Two or more foreground tasks running simultaneously is when context switching takes a real toll. I keep it to one.&lt;/p&gt;
&lt;h3 id="background-light-supervision"&gt;Background: light supervision&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;One or two background tasks. Check on them when you have a gap in the foreground task.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;While I write this post, I have an autolearning plugin and skill system running in the background. Useful work, but no urgent priority. I started the agent on a task before I sat down to write, and it works while I focus on prose. When I take a break, I check in, review what it did, give feedback, and go back to writing.&lt;/p&gt;
&lt;p&gt;That is the background pattern: you started the task, you trust it, and you only check when there is a natural gap. Maybe the foreground agent is thinking, or running a test suite. You switch over, do a light-touch review, and switch back. The harness handles the work; you just need to know when it is done.&lt;/p&gt;
&lt;p&gt;One or two background tasks is the sweet spot. It's tempting to add more; I know because I tried! I can handle two. On occasion I can handle three. Four will regularly push me past my limit.&lt;/p&gt;
&lt;h3 id="shadow-automations-zero-daily-attention"&gt;Shadow automations: zero daily attention&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;As many as you want. They run on a schedule and enrich your context without asking anything of you.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Beyond foreground and background, there is a third tier: fully automated tasks that run on a schedule and need zero interactive supervision. I think of these as automations running in the shadow. They run in the background of the background, tending things so you do not have to.&lt;/p&gt;
&lt;p&gt;I have an automation that periodically scours my knowledge vault and links notes together that should be connected. I never do this manually; the agent handles it entirely. Another one monitors my GitHub repos for merge-able Dependabot PRs and merges them on its own. A third scans my vault every hour for files with today's date in &lt;a href="https://en.wikipedia.org/wiki/ISO_8601"&gt;ISO 8601&lt;/a&gt; format and enriches them through a state machine: transcript becomes meeting note, meeting note becomes daily bullet. I used to do this manually, copying and pasting meeting notes into my knowledge base and prompting the same thing over and over. It was mostly tedious and sometimes error-prone. Once I automated it, maintaining the vault became much easier.&lt;/p&gt;
&lt;p&gt;One more: at work, an automation runs at 4:30 pm every day and gives me a summary of my GitHub activity. Commits, pull requests, reviews, issues. I get a digest without lifting a finger.&lt;/p&gt;
&lt;p&gt;You only check on these at the end of the day. Did they run? Did they do the right thing? A quick review, and you are done. This tier is where natural-language automation shines. In the &lt;a href="https://openai.com/codex/"&gt;Codex&lt;/a&gt; app, you can describe the task, tell it when to run, and even tell it to delete itself when it is finished. Ephemeral automations are self-cleaning: once the task is done, the agent knows to delete the automation, so you never accumulate stale scheduled jobs.&lt;/p&gt;
&lt;h2 id="you-still-need-a-priority-list"&gt;You still need a priority list&lt;/h2&gt;&lt;p&gt;The three tiers tell you &lt;em&gt;how&lt;/em&gt; to run tasks. They do not tell you &lt;em&gt;what&lt;/em&gt; to run. For that, you need a plain priority list: a task list for the day, ordered by what matters most.&lt;/p&gt;
&lt;p&gt;Nothing fancy. I write down the things I want to get done, rank them, and that becomes the input to the three-tier system. The top item goes to foreground. The next couple go to background. Shadow automations run regardless, because they do not compete for my attention.&lt;/p&gt;
&lt;p&gt;The goal is straightforward: if you knock off your major tasks for the day, the system is working. The tiers are there to help you do that without juggling everything at once.&lt;/p&gt;
&lt;h2 id="how-notifications-tie-it-together"&gt;How notifications tie it together&lt;/h2&gt;&lt;p&gt;Notifications turn "I hope the background task is fine" into "I will know when it needs me." Without a notification system, you have to manually check tabs and windows, which is exactly the kind of anxious context switching the three-tier approach is designed to eliminate.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://openai.com/codex/"&gt;Codex&lt;/a&gt;, &lt;a href="https://github.com/manaflow-ai/cmux"&gt;&lt;code&gt;cmux&lt;/code&gt;&lt;/a&gt;, and &lt;a href="https://www.cursor.com"&gt;Cursor&lt;/a&gt; all provide notification indicators -- dots, badges, rings, color changes. You glance at the sidebar and know what needs attention. The foreground task keeps running. Your focus stays intact.&lt;/p&gt;
&lt;p&gt;When a notification arrives, you decide: does this need me now, or can it wait? If it can wait, let it. The background task is lower priority by definition.&lt;/p&gt;
&lt;h2 id="the-framework-in-practice"&gt;The framework in practice&lt;/h2&gt;&lt;p&gt;Right now, as I finish this post, my three tiers look like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Foreground&lt;/strong&gt;: writing this blog post, iterating on prose and voice with my agent.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Background&lt;/strong&gt;: developing my autolearning plugin and skill system, checked between writing sessions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shadow automations&lt;/strong&gt;: vault linking, Dependabot merging, daily GitHub digests, meeting note enrichment. All running on their own.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The harness runs each task, tracks its state, and pings you when something needs attention. You stop keeping tabs in your head and decide only what deserves your foreground focus.&lt;/p&gt;
&lt;h2 id="move-things-down-the-chain"&gt;Move things down the chain&lt;/h2&gt;&lt;p&gt;The three-tier framework is not static. The real leverage comes from constantly asking: can I push this further down the chain? Can this foreground task become a background task? Can this background task become a shadow automation?&lt;/p&gt;
&lt;p&gt;The more you move tasks from foreground to background to silent automation, the more effortless your multitasking becomes. You are removing yourself from the loop on tasks that do not need you.&lt;/p&gt;
&lt;p&gt;That said, be pragmatic. Some tasks genuinely need your judgment, and over-automating creates brittle systems that break in ways you do not notice. Push things down the chain when it makes sense, not just because you can.&lt;/p&gt;
&lt;p&gt;Pick a harness. Try the three tiers for a week. See how it feels.&lt;/p&gt;
</content></entry><entry><title>Exploring Agent Harnesses - A Comparative Analysis of Codex, CMux, and Cursor</title><link href="https://ericmjl.github.io/blog/2026/6/4/exploring-agent-harnesses/" rel="alternate"/><updated>2026-06-04T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:babc5fcc-b1d6-3b67-a046-a876ed22d595</id><content type="html">&lt;p&gt;I have used three coding agent harnesses extensively over the past few months: the &lt;a href="https://openai.com/codex/"&gt;Codex&lt;/a&gt; app, &lt;a href="https://github.com/manaflow-ai/cmux"&gt;&lt;code&gt;cmux&lt;/code&gt;&lt;/a&gt;, and &lt;a href="https://www.cursor.com"&gt;Cursor&lt;/a&gt;. All three have been immense in amplifying my productivity. All three share a similar layout: workspaces on the left, a coding agent in the middle, and a terminal on the right. A terminal and a browser are really all you need. I decided to write up how they compare.&lt;/p&gt;
&lt;h2 id="the-comparison-matrix"&gt;The comparison matrix&lt;/h2&gt;&lt;p&gt;I decided to compare them against the following axes:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;&lt;code&gt;cmux&lt;/code&gt; + OpenCode&lt;/th&gt;
&lt;th&gt;Cursor&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Workspace view&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Partial&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Notification system&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Natural-language automations&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open source&lt;/td&gt;
&lt;td&gt;Yes (Apache-2.0 CLI)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Built-in browser&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes (element selection)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;ul&gt;
&lt;li&gt;Workspace view refers to whether the harness provides a visual workspace abstraction where one workspace corresponds to one project, and each workspace contains multiple chat threads.&lt;/li&gt;
&lt;li&gt;Notification system is whether the harness provides visual indicators (dots, badges, rings) that show which threads or workspaces need your attention.&lt;/li&gt;
&lt;li&gt;Natural-language automations is whether you can set up recurring or ephemeral tasks by describing them in plain language rather than editing &lt;code&gt;cron&lt;/code&gt; or writing scripts by hand.&lt;/li&gt;
&lt;li&gt;Open source is obvious: whether it's open source or not.&lt;/li&gt;
&lt;li&gt;Built-in browser is whether the harness ships with a browser pane so you can preview web apps without leaving the tool.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With those defined, let's dive in.&lt;/p&gt;
&lt;h2 id="workspace-management"&gt;Workspace management&lt;/h2&gt;&lt;p&gt;Before we go on, I think it'll be helpful to disambiguate some terms. In the following paragraphs, when we talk about a workspace, it really refers to one project, or one folder, or one git repo. (Incidentally, this is in line with the &lt;a href="https://ericmjl.github.io/data-science-bootstrap-notes/project/"&gt;1:1:1:1... rule&lt;/a&gt; that I advocate for in the Data Science Bootstrap Notes.) Each workspace contains multiple chat threads, one per task or conversation with the agent. You might have one workspace for your web app, another for a data pipeline, a third for a side project. Within each workspace, you spin up threads as needed: one to debug a failing test, another to add a feature, another to review a pull request.&lt;/p&gt;
&lt;p&gt;Effectively, workspaces group related work, while threads keep individual tasks isolated. As a coder, you will want to see all of this at a glance, switch between threads, and know which ones are active. Without workspaces, you are stuck juggling unrelated conversations in a flat list, constantly losing context when you switch tasks.&lt;/p&gt;
&lt;p&gt;Codex and Cursor both handle this well. You can see a list of workspaces on the left sidebar, expand into threads within each one, and pick up where you left off. Workspaces group your projects, threads track your conversations. This two-level model, workspaces plus threads, is the right visual abstraction for agent harnesses. Cursor's UI looks like the following:&lt;/p&gt;
&lt;p&gt;&lt;img src="cursor-workspaces.webp" alt="Cursor workspace UI showing multiple workspaces on the left sidebar"&gt;&lt;/p&gt;
&lt;p&gt;On the other hand Codex's UI looks like the following:&lt;/p&gt;
&lt;p&gt;&lt;img src="codex-workspaces.webp" alt="Codex workspace UI showing workspaces on the left sidebar"&gt;&lt;/p&gt;
&lt;p&gt;The commonality is that workspaces (i.e. projects) are on the left, and I can quickly switch between them.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;cmux&lt;/code&gt; provides workspace-level tabs. When paired with OpenCode, threads live either as individual tabs within a &lt;code&gt;cmux&lt;/code&gt; workspace (more visually obvious) or as sessions within OpenCode itself (less visually obvious, since they are not visually separated). Visually, it looks like this:&lt;/p&gt;
&lt;p&gt;&lt;img src="cmux-workspaces.webp" alt="cmux workspace UI showing vertical tabs"&gt;&lt;/p&gt;
&lt;h2 id="notification-systems"&gt;Notification systems&lt;/h2&gt;&lt;p&gt;Workspaces tell you what is running. Notifications tell you when something finishes or needs attention.&lt;/p&gt;
&lt;p&gt;Codex and Cursor both show clear indicators: a dot, a badge, or a color change on the workspace tab when an agent completes a task. You can glance at the sidebar and know what is done without scrolling through tabs.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;cmux&lt;/code&gt; has notification features: blue rings on tabs and a sidebar notification panel. When I first started using it, notifications were not yet exposed. I submitted a pull request to a fork that was working on them, and the feature has since shipped in &lt;code&gt;cmux&lt;/code&gt;. You can see at a glance which tasks have finished.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;cmux&lt;/code&gt; supports multiple agents including OpenCode, with hooks for session restore and notification wiring. Without notifications, you have to manually scroll through tabs and windows to check on background tasks, which defeats the purpose of multitasking. Notifications are only the second layer. The third piece is automation.&lt;/p&gt;
&lt;h2 id="automations"&gt;Automations&lt;/h2&gt;&lt;p&gt;Automations are where things get exciting. All three harnesses let you set up recurring tasks, but they differ in how you create them and how you manage them.&lt;/p&gt;
&lt;p&gt;Codex allows you to create automations using natural language from within a workspace. You can tell it: "Set up an automation that checks GitHub every five minutes for a deployment to finish, and then delete the automation when it is done." These ephemeral automations are self-cleaning. You describe the task, tell it to shut itself down when complete, and Codex handles the rest.&lt;/p&gt;
&lt;p&gt;Cursor supports similar automation workflows. You can configure recurring tasks and let them run in the background.&lt;/p&gt;
&lt;p&gt;The OpenCode ecosystem has scheduling capabilities through plugins and community tools. You configure automations using natural language, and the plugin translates that into a scheduled task. The downside is that there is no unified user interface for viewing or managing these automations yet. You can set them up, but you cannot easily see what is running (yet). That is another gap worth noting, but knowing the open source world, it'll get better!&lt;/p&gt;
&lt;h2 id="built-in-browser"&gt;Built-in browser&lt;/h2&gt;&lt;p&gt;All three come with a built-in browser, which is a big deal if you are building web apps. You can see your app running alongside the agent, iterate on changes, and verify results without leaving the harness.&lt;/p&gt;
&lt;p&gt;&lt;img src="cursor-browser.webp" alt="Cursor&amp;#39;s built-in browser showing a web app alongside the coding agent"&gt;&lt;/p&gt;
&lt;p&gt;Cursor is the standout here. It lets you select individual DOM elements and add them directly into the context window. When you are building a web app and want the agent to focus on a specific component, you can point at it and say "fix this" without having to describe where it is. That kind of selective context control makes a real difference for front-end work.&lt;/p&gt;
&lt;p&gt;&lt;img src="cursor-element-selection.webp" alt="Cursor element selection feature highlighting a DOM element"&gt;&lt;/p&gt;
&lt;h2 id="open-source-status"&gt;Open source status&lt;/h2&gt;&lt;p&gt;OpenCode, &lt;code&gt;cmux&lt;/code&gt;, and the &lt;a href="https://github.com/openai/codex"&gt;Codex CLI&lt;/a&gt; (Apache-2.0) are open source. Cursor is proprietary. The Codex desktop app builds on the open-source CLI but is itself a closed product.&lt;/p&gt;
&lt;p&gt;This matters more to some people than others. I grew up in the open-source world during grad school, and I gravitate toward open tools. With open-source tools, if something annoys you, you can fix it yourself. In the age of agents, fixing tool limitations is more a matter of imagination than skill.&lt;/p&gt;
&lt;p&gt;That said, I am not going to be pushy about open source here. It is a personal preference; use what works for you.&lt;/p&gt;
&lt;h2 id="choosing-a-harness"&gt;Choosing a harness&lt;/h2&gt;&lt;p&gt;All three harnesses I've gone through here are pretty capable. I obviously did not cover Claude Code, it's well-known by now; these are just the three that I've used. They're all pretty capable; once set up properly, they handle roughly the same range of tasks.&lt;/p&gt;
&lt;p&gt;For me, the deciding factors are practical. If your company mandates a specific tool, use that at work. &lt;strong&gt;Then pick a different one for home.&lt;/strong&gt; Visual separation between work and home environments matters more than you might think. I &lt;a href="https://ericmjl.github.io/blog/2026/4/28/how-i-recognized-and-handled-ai-burnout/"&gt;burned out in April&lt;/a&gt; partly because I was using &lt;code&gt;cmux&lt;/code&gt; with OpenCode at both work and home -- the same interface, even on different laptops, with no visual disconnect. Effectively, my brain never left the office.&lt;/p&gt;
&lt;p&gt;Using different tools at work and at home has a bonus side effect: you avoid muscle-memory lock-in to a single vendor. You stay flexible.&lt;/p&gt;
&lt;p&gt;Cost is another factor worth considering. If cost is a factor, OpenCode plus &lt;code&gt;cmux&lt;/code&gt; with open-weight models through providers like &lt;a href="https://openrouter.ai"&gt;OpenRouter&lt;/a&gt; can be significantly cheaper than subscription-based options. The free tier on some model providers goes a long way.&lt;/p&gt;
&lt;h2 id="try-them-for-yourself"&gt;Try them for yourself&lt;/h2&gt;&lt;p&gt;If you're looking for counsel on what to pick, this is what I would say: pick one, use it for a few weeks to build real familiarity, then try another. Give yourself enough time with each one to form an honest opinion. These tools are evolving fast, and the landscape will look different six months from now! Only by trying things out can you know what suits you best.&lt;/p&gt;
</content></entry><entry><title>Reflections from the BioIT World workshop - standardization is worth the effort</title><link href="https://ericmjl.github.io/blog/2026/5/27/reflections-from-bioit-world-workshop-standardization-is-worth-the-effort/" rel="alternate"/><updated>2026-05-27T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:6d82c16f-9ef5-3e5f-aa95-97e4ad9e3a3c</id><content type="html">&lt;p&gt;My teammate &lt;a href="https://jackievaleri.github.io/"&gt;Jackie Valeri&lt;/a&gt; and I recently co-taught a workshop at &lt;a href="https://www.bio-itworldexpo.com/"&gt;BioIT World 2026&lt;/a&gt; on standardizing data science ways of working. Walking out of the room, I kept thinking about the same tension I have seen for years: most teams already feel the pain of missing standards. They have a harder time believing the effort pays off, and a harder time doing the people work required to make change stick.&lt;/p&gt;
&lt;p&gt;That tension is what this post is about. It is my reflection on what we covered, what landed, and what I wish we had emphasized more. The through-line is simple: standardization is worth the effort. The sections below walk through why, how to choose what to standardize, how to get buy-in, and how to keep standards alive as the stack changes.&lt;/p&gt;
&lt;h2 id="pain-comes-before-payoff"&gt;Pain comes before payoff&lt;/h2&gt;&lt;p&gt;Most teams need to feel the pain before they believe the payoff. Two stories from my time at Moderna show both sides.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vignette 1: the messy codebase.&lt;/strong&gt; Early in my time at Moderna, I studied older codebases that had no shared structure. No standard folder layout, no predictable onboarding command, no consistent way to find tests or documentation. I pitched &lt;a href="https://www.linkedin.com/in/giessel/"&gt;Andrew Giessel&lt;/a&gt; that we had to standardize while the team was still small: if we kept going like this, onboarding would stay painful, and when the bus factor hit, we would feel it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vignette 2: the easy migrations.&lt;/strong&gt; Later, when we migrated from Bitbucket to GitHub, and from Conda to &lt;a href="https://pixi.sh/"&gt;Pixi&lt;/a&gt;, the logistics were surprisingly smooth. We had one pattern to upgrade, one CLI command to run, and minimal manual tweaking per project. Part of that was team size (we are roughly twenty data scientists + a few dozen more adjacent folks, not a few hundred). Part of it was the fruit of previously hiring people willing to learn the stack (more on that below). The biggest part was standardization itself: we knew exactly what we were migrating because we had standardized it in the first place.&lt;/p&gt;
&lt;p&gt;The contrast is the point. Vignette 1 is what happens without standards; vignette 2 is what opens up once you have them. The upfront cost buys confidence, predictability, and speed later.&lt;/p&gt;
&lt;h2 id="work-backwards-from-what-you-ship"&gt;Work backwards from what you ship&lt;/h2&gt;&lt;p&gt;Believing the payoff exists is step one. Step two is choosing which standards actually matter.&lt;/p&gt;
&lt;p&gt;At Moderna, we applied step two by naming our deliverables first:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Compute tasks&lt;/strong&gt;: CLI tools we can run in the cloud with as much scalability as we can manage.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Python packages&lt;/strong&gt;: reusable components other computational scientists can import and build on.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Once we named those artifacts, the required standards became obvious. Compute tasks need Dockerfiles (containers) and CI/CD (deployment pipelines) that let us move quickly. Python packages need consistent project scaffolding, tests, and documentation.&lt;/p&gt;
&lt;p&gt;That backward reasoning is how we chose which standards to invest in. I learned that lesson the hard way in grad school, where bad software patterns cost me days of retesting on the HPC because my code was organized for batch runs, not incremental checks. I have also heard familiar horror stories from finance, where a one-month notebook prototype became an eight-month production slog after handoff to an engineering team on a different stack. I knew which pain points were real because I had felt some of them directly and heard about others often enough.&lt;/p&gt;
&lt;p&gt;Deliverables will likely differ, but the principle stays the same: name the artifact, and ask which practices remove friction on the path to shipping it.&lt;/p&gt;
&lt;h2 id="standardize-where-divergence-creates-friction"&gt;Standardize where divergence creates friction&lt;/h2&gt;&lt;p&gt;Naming your artifacts tells you the destination. It does not tell you where to invest first. Day-to-day friction does.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Standardize the places where inconsistent approaches slow the team down.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For us, the clearest example was project template structure. We had code in Julia, R, and Python. Even within Python, different projects used different frameworks. Onboarding meant learning a new layout every time. There was no single command to get started.&lt;/p&gt;
&lt;p&gt;Once we standardized scaffolding, onboarding got faster. Jumping into a colleague's repo became predictable: I know where the docs live, where the tests live, and how to run things. That predictability compounds.&lt;/p&gt;
&lt;p&gt;That does not mean we standardize everything. Algorithm choices, analysis approaches, and visualization design stay flexible. We focus standards on interfaces and workflows, the parts that help people collaborate without constraining the science.&lt;/p&gt;
&lt;p&gt;Even with that selectivity, one pattern is reliable: if your team keeps re-explaining the same basics, that is where to start. Folder structure is a common answer. Documentation publishing is another strong first move when everyone already knows docs matter but nobody has a single place to put them.&lt;/p&gt;
&lt;h2 id="people-problems-beat-roi-math"&gt;People problems beat ROI math&lt;/h2&gt;&lt;p&gt;Knowing where to standardize is the easy part. Getting people to adopt it is harder.&lt;/p&gt;
&lt;p&gt;Adoption is a people problem, but you still have to get in the room first. That is why we spent real time in the workshop on ROI. Moderna has a strong ROI culture; working with &lt;a href="https://www.linkedin.com/in/dave-johnson-60a95142/"&gt;Dave Johnson&lt;/a&gt; early on reinforced that return on investment matters when you pitch an initiative. We shared resources on how to calculate it, and we walked through the concept: build the case before you standardize, as the argument you bring to the table. (For a concrete example, I once worked through the math for our quarterly docathons in &lt;a href="../../../../2024/6/30/two-years-of-docathons-insights-and-lessons-learned/"&gt;Two years of docathons: Insights and lessons learned&lt;/a&gt;.) The ROI frame of mind helps with managers who think in business terms; use it to make the strongest case you can.&lt;/p&gt;
&lt;p&gt;A strong ROI case opens the door. It does not carry people through the migration. We are nerds, and we would rather debug a conda solve than navigate disagreement about tooling. Technology helps, but the people work still carries the change. Getting buy-in, training people, and holding hands through a migration is still work. At the end of the day, "hand-holding" is training, repeated until the new pattern feels normal.&lt;/p&gt;
&lt;p&gt;I learned that split the hard way when I pitched standardization to Andrew Giessel years ago. I leaned on open-source patterns I knew would be cost-effective. The tooling was only half the pitch. The other half was training people on those toolsets and building a culture that maintains them together.&lt;/p&gt;
&lt;p&gt;If you take one people tactic from this post, make it this: use ROI to open the door, then budget time for training and socialization after the demo. Adoption is a milestone, not a finish line.&lt;/p&gt;
&lt;h2 id="standards-evolve"&gt;Standards evolve&lt;/h2&gt;&lt;p&gt;Training gets people onto a standard. It does not freeze the standard in place. The technology landscape keeps moving, so your standards have to move with it.&lt;/p&gt;
&lt;p&gt;BioIT World gave us a live example of that drift. &lt;a href="https://marimo.io/"&gt;Marimo&lt;/a&gt; notebooks were part of our tech demo, and the audience reacted differently to Marimo than to our CLI tooling. Marimo looked cool, sure, but the deeper reaction was: "This is mindblowing! I had no idea we could interact with notebooks this way!" It was a glimpse of something new.&lt;/p&gt;
&lt;p&gt;The excitement was real. Whether, and when, to move to Marimo notebooks is still an open decision. We are in the middle of that decision now. Jupyter notebooks served us for years. Marimo offers real advantages, especially because coding agents can work with Marimo notebooks more controllably than with classic Jupyter. I wrote more about that pattern in &lt;a href="../../../../2025/10/28/use-coding-agents-to-write-marimo-notebooks/"&gt;Use coding agents to write Marimo notebooks&lt;/a&gt;. Marimo is also Python-only today, which sits awkwardly next to our multi-language history. That tension is normal.&lt;/p&gt;
&lt;p&gt;While we decide, the playbook stays the same: demo the idea, get review from peers, socialize the benefits, then train people through the migration. We have done this before with GitHub, Pixi, and other stack changes. Standardization made each migration tractable because we only had one pattern to evolve. Pixi is one example of how we choose tools; Conda may be the right call in your environment. The principle matters more than the brand name. Stay close enough operationally that upgrades stay manageable.&lt;/p&gt;
&lt;h2 id="ai-lowers-the-cost"&gt;AI lowers the cost&lt;/h2&gt;&lt;p&gt;As standards keep evolving, the cost of building and migrating tooling keeps dropping. That is the shift I want to name in this section.&lt;/p&gt;
&lt;p&gt;Building the CLI commands that made our GitHub and Pixi migrations tractable used to be the hard part. Today you can describe what you want to a coding agent and iterate until you have a &lt;a href="https://github.com/ericmjl/pyds-cli"&gt;&lt;code&gt;pyds-cli&lt;/code&gt;&lt;/a&gt;-style scaffold generator, a migration helper, or a project bootstrapper. The effort is lower; the ROI is higher.&lt;/p&gt;
&lt;p&gt;I was already experimenting in this direction during those migrations. My first LLM-side experiment was a git commit message writer. The one that mattered for the migrations was a CLI helper that used models to propose intelligent file merges, with a human still reviewing the result. Same shape as the tools I am describing now; the difference is how quickly you can build them.&lt;/p&gt;
&lt;p&gt;Cheaper tooling is only half the story. AI also amplifies whatever patterns already exist in your codebase. Feed it good standards, and it follows them. Leave chaos in the repo, and it reproduces chaos faster. The tools got cheaper; the human coordination problem still needs attention. Standardization pays off on both sides.&lt;/p&gt;
&lt;h2 id="hiring-matters-too"&gt;Hiring matters too&lt;/h2&gt;&lt;p&gt;That coordination problem is where hiring and culture matter. I flagged this in vignette 2, and it is worth stating plainly: easy migrations are a people win as much as a tooling win.&lt;/p&gt;
&lt;p&gt;Standardization works best when people are willing to learn new parts of the stack. Software skill levels can vary widely. The non-negotiable trait, for us, is curiosity about the tooling we use to do science. That is a form of gatekeeping, and I am fine saying so. If someone refuses to adapt when the team agrees on a shared pattern, friction returns.&lt;/p&gt;
&lt;p&gt;Hiring for learning agility is one place leaders can start, especially if they are building a team from scratch. If you already have a team in place, you cannot rehire your way out of friction; start with the software pattern that removes the most daily friction, then grow the culture from there.&lt;/p&gt;
&lt;h2 id="what-would-you-start-with"&gt;What would you start with?&lt;/h2&gt;&lt;p&gt;You have seen the payoff, the decision rules, the people work, and the stack as it keeps moving. The remaining question is practical: where would you start?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is one thing you would like to start standardizing for your team?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Pick one. Folder structure. Dependency management. CI/CD for every repo. Documentation templates. One choice, executed well, beats a grand roadmap that rarely ships.&lt;/p&gt;
&lt;p&gt;Getting started can be as small as a half-day survey of tools in your landscape, or a conversation with your manager backed by an ROI estimate. You can even ask an AI assistant to walk you through Pixi, &lt;a href="https://cookiecutter.readthedocs.io/en/latest/"&gt;Cookiecutter&lt;/a&gt;, or whatever tool you are evaluating. The specific path is yours; the important part is to start.&lt;/p&gt;
&lt;p&gt;If you want to go deeper before you pick that one thing, two resources cover adjacent ground. For delivery models, scaffolding, and implementation tactics, see my blog post, &lt;a href="../../../../2025/4/2/how-to-standardize-data-science-ways-of-working-to-unlock-your-teams-creativity/"&gt;How to Standardize Data Science Ways of Working to Unlock Your Team's Creativity&lt;/a&gt;. For machine setup, project structure, and the fundamentals underneath all of this, see my online eBook, &lt;a href="https://ericmjl.github.io/data-science-bootstrap-notes/"&gt;The Data Science Bootstrap Notes&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="acknowledgments"&gt;Acknowledgments&lt;/h2&gt;&lt;p&gt;This work has always been a team sport. With thanks to &lt;a href="https://www.linkedin.com/in/giessel/"&gt;Andrew Giessel&lt;/a&gt;, &lt;a href="https://www.linkedin.com/in/dave-johnson-60a95142/"&gt;Dave Johnson&lt;/a&gt;, &lt;a href="https://www.linkedin.com/in/adriannaloback/"&gt;Adrianna Loback&lt;/a&gt;, &lt;a href="https://www.linkedin.com/in/rlvislaywadephd"&gt;Rebecca Vislay-Wade&lt;/a&gt;, &lt;a href="https://jackievaleri.github.io/"&gt;Jackie Valeri&lt;/a&gt;, &lt;a href="https://www.linkedin.com/in/albert-lam/"&gt;Albert Lam&lt;/a&gt;, Anand Murthy, &lt;a href="https://www.linkedin.com/in/dandluu/"&gt;Dan Luu&lt;/a&gt;, and other colleagues who helped design, maintain, and evolve our standards over the years. Andrew and Dave gave us the freedom to build; my current manager &lt;a href="https://www.linkedin.com/in/jwadedavis"&gt;Wade Davis&lt;/a&gt; continues to give us operational room to make it happen, and for that I am grateful!&lt;/p&gt;
</content></entry><entry><title>What data science is actually about in the age of AI</title><link href="https://ericmjl.github.io/blog/2026/5/20/what-data-science-is-actually-about-in-the-age-of-ai/" rel="alternate"/><updated>2026-05-20T00:00:00Z</updated><author><name>Eric Ma</name></author><id>urn:uuid:9fc693be-c389-3139-8ff9-4e4bda1632d1</id><content type="html">&lt;p&gt;I have been thinking about what the core mission of a data scientist actually is in 2026, surrounded by AI coding assistants, LLM-powered applications, and the relentless buzz around adaptive software development. The answer keeps coming back the same: measurement.&lt;/p&gt;
&lt;p&gt;LLMs and AI coding tools are new instruments for that work, not replacements for it. And instruments do not decide what to measure. People with domain knowledge do. The principle is simple and ought to be stated plainly: measurement must be defined by the people closest to the problem.&lt;/p&gt;
&lt;p&gt;But staying close to the problem is harder than it sounds. Something is pulling data scientists away from that mission, toward a very different self-image. It worries me.&lt;/p&gt;
&lt;h2 id="the-seduction-of-full-stack"&gt;The seduction of full-stack&lt;/h2&gt;&lt;p&gt;I keep running into a narrative that data scientists should become full-stack software developers. The logic goes something like this: you can already code, and now AI makes coding even easier, so why are you not contributing to the production codebase?&lt;/p&gt;
&lt;p&gt;Managers hear that AI-assisted coding has arrived and think every data scientist on their team can now own a slice of the front-end, the back-end, the database, and the deployment pipeline. Data scientists get pulled into writing TypeScript for UI components, setting up production databases, and building integrations with other systems, because those tasks fill the week. The work that made them valuable in the first place (defining metrics, designing evaluation frameworks, rigorously measuring whether a system actually works) gets squeezed into the margins.&lt;/p&gt;
&lt;p&gt;This is not an accident or an outlier. The whole industry is zigging toward turning data scientists into app developers (&lt;a href="https://www.acceldata.io/blog/convergence-of-personas-how-ai-is-reshaping-data-management-functions"&gt;Acceldata&lt;/a&gt;, &lt;a href="https://www.deloitte.com/us/en/Industries/tmt/articles/ai-software-development-engineering-roles-being-rewritten.html"&gt;Deloitte&lt;/a&gt;, &lt;a href="https://www.linkedin.com/pulse/full-stack-data-scientist-responsibilities-across-lifecycle-cry6f"&gt;GSDC&lt;/a&gt;). The zig is understandable: building a GenAI system has never been easier. Call an API, wire up a prompt, ship a demo. But the easy part is not where the gap is. The gap is reliability, ensuring these systems actually work in production, at scale, over time. That is the zag, and it is where data scientists create massive alpha.&lt;/p&gt;
&lt;h2 id="why-this-is-a-mistake"&gt;Why this is a mistake&lt;/h2&gt;&lt;p&gt;The zig is tempting because building is seductive. But AI amplifies strengths and weaknesses alike, and a team of generalists building sloppily gets amplified into a production incident. Your ability to zag depends on having people trained for reliability work, and pulling them into the build instead defeats the purpose. What data scientists contribute is not a stack of engineering skills but a form of reasoning that takes years to develop, and that reasoning is exactly what GenAI systems need most.&lt;/p&gt;
&lt;p&gt;Most data scientists I know came from scientific backgrounds. Economists, biologists, chemists, sociologists. Their prior training is in hypothesis generation, experimental design, measurement, and rigorous analysis. That training is hard to replicate. It takes years to move from knowing the techniques to developing the scientific judgment needed to look at a set of results and know whether they hold up.&lt;/p&gt;
&lt;p&gt;When you pull a data scientist into full-stack software development, you are trading that hard-won scientific expertise for a second-rate software engineer. Data scientists will make immature decisions about database selection, front-end framework choices, and system architecture that an experienced software developer would avoid. The engineering suffers from inexperience, and the science suffers from neglect.&lt;/p&gt;
&lt;p&gt;The cost shows up in two ways. First, cognitive switching between measurement thinking and software engineering thinking degrades both. Building a UI component requires a very different mode of thought than designing an evaluation framework for a stochastic system. Second, the measurement work itself gets displaced entirely. The data scientist who should be spending their week figuring out whether the LLM pipeline is actually performing well is instead debugging a CSS layout.&lt;/p&gt;
&lt;p&gt;The framing I want every data scientist and their manager to internalize is a trade-off. Every hour spent maintaining databases and front-ends is an hour that cannot be spent on the measurement work the team is hired to do.&lt;/p&gt;
&lt;h2 id="what-data-scientists-should-focus-on"&gt;What data scientists should focus on&lt;/h2&gt;&lt;p&gt;So if full-stack development is the wrong direction, what is the right one?&lt;/p&gt;
&lt;p&gt;The clearest answer is LLM evaluation. Every company building with LLMs needs to know whether their system works. The answer requires exactly the kind of rigorous, scientifically grounded measurement that data scientists are trained to do.&lt;/p&gt;
&lt;p&gt;The process is straightforward. You need labeled data that represents what "good" looks like. You run your model against that data. You write programmatic checks to evaluate the model's outputs. You quantify metrics based on those checks. Then you iterate.&lt;/p&gt;
&lt;p&gt;The tools for doing this are immature, which is actually an opportunity. You do not need an expensive vendor evaluation platform. What you need is the ability to select and label your own data, store it in a simple database, and run checks against it. This is where data scientists can and should build utilities for their own work. I will come back to the line between building a measurement utility and practicing software engineering shortly.&lt;/p&gt;
&lt;p&gt;Building the utilities is the straightforward part. The harder challenge is the scientific logic. Deciding what to measure, debating which metric actually reflects the business problem, and figuring out whether an improvement is real or noise all require exactly the thinking trained in graduate programs. Hamel Husain makes this concrete in his post on &lt;a href="https://hamel.dev/blog/posts/revenge/"&gt;the revenge of the data scientist&lt;/a&gt;: every common eval pitfall, from generic metrics to unverified judges to bad experimental design, traces back to a missing data science fundamental.&lt;/p&gt;
&lt;h2 id="where-tool-building-ends-and-software-engineering-begins"&gt;Where tool-building ends and software engineering begins&lt;/h2&gt;&lt;p&gt;The line between building a measurement utility and practicing software engineering is the nuance that matters most, so let me be specific.&lt;/p&gt;
&lt;p&gt;Data scientists should absolutely build tools for their own measurement work. The principle is simple: build single-purpose, tightly scoped utilities with no feature creep, where you are the primary user. I have built CLI tools that do one thing and do it well, in the Unix tradition. I built them for myself first, and if someone else happens to have the exact same problem, they can use the tool as-is. That is the dogfooding principle in action.&lt;/p&gt;
&lt;p&gt;For example, if I need to do LLM evaluations, I need labeled data. Instead of waiting for a software development team to build a full-stack application for data labeling, I can quickly set up a lightweight database with a simple UI that lets collaborators label data. The whole setup is a measurement tool, not a product.&lt;/p&gt;
&lt;p&gt;The line between tool and product turns on three criteria.&lt;/p&gt;
&lt;p&gt;The first is scope. If the tool does one high-value thing really well, it belongs in the data scientist's remit to build. If it needs to evolve to handle diverse use cases, feature requests, and edge cases from dozens or hundreds of users, it has crossed into software engineering territory.&lt;/p&gt;
&lt;p&gt;The second is audience. If you are the primary user and others incidentally benefit, you are building a measurement tool. If you are building something primarily for other people to use, you are doing product development, and that belongs with the software engineering team.&lt;/p&gt;
&lt;p&gt;The third is the production boundary. It is fine for data scientists to prototype UIs to show what is possible, and to own the model and its API endpoint. It is fine to build tooling that supports data collection and measurement. But the finished product belongs with professional software engineers: the chat system integrations, the production database, the user-facing front-end, the infrastructure that other teams depend on.&lt;/p&gt;
&lt;p&gt;When you cross that line, data engineers (the people who build and maintain production data pipelines) and software engineers should own the systems. Data scientists can be team players and pitch in when needed, but maintaining those systems should never become their full-time responsibility.&lt;/p&gt;
&lt;h2 id="measurement-done-right-in-practice"&gt;Measurement done right, in practice&lt;/h2&gt;&lt;p&gt;Ownership boundaries are one kind of clarity. Another is knowing how to measure well. Let me make this concrete with two examples from different corners of data science.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Document parsing for information retrieval.&lt;/strong&gt; This is a problem nearly every team working with LLMs faces. You have a pile of documents and you need to extract structured information from them. The natural metrics are precision and recall. But here is the trap: if you optimize only for recall, you extract too much noise, and someone downstream has to filter garbage. If you optimize only for precision, you miss important information that actually matters for decisions downstream.&lt;/p&gt;
&lt;p&gt;Picking one metric in isolation leads to real problems. The right measurement framework has to be debated and co-defined with the people who understand what the extracted information will be used for. The data scientist brings the rigor; the domain experts bring the context. Neither can do it alone.&lt;/p&gt;
&lt;p&gt;The same pattern appears in a completely different setting. &lt;strong&gt;ML models for property prediction in preclinical research.&lt;/strong&gt; A data science team builds a machine learning model to predict molecular properties. The model has an R-squared value and other statistical metrics. But those numbers have to translate into real decisions: which mutant should we select for the next round of engineering, which compound is worth synthesizing.&lt;/p&gt;
&lt;p&gt;The metrics that matter are co-defined between the lab team and the data science team. The lab scientists know what actually matters for their experiments, and the data scientists know how to build models and measure their performance. Together, they arrive at a measurement framework that bridges statistical performance and experimental reality.&lt;/p&gt;
&lt;p&gt;Both examples share a common pattern. The measurement framework is designed by the people closest to the problem, not imposed from above or derived from whatever metric is easiest to compute.&lt;/p&gt;
&lt;h2 id="for-managers-reframing-the-conversation"&gt;For managers - reframing the conversation&lt;/h2&gt;&lt;p&gt;If you manage a data science team, or if you are an executive making staffing decisions, the question is: what should you actually do with this framing? Here are four practices I want you to take away.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reframe how you define problems.&lt;/strong&gt; Instead of framing work as "we are going to build a solution," frame it as "we are going to figure out how good this solution is." Measurement comes first, not after. Embed it throughout the product development lifecycle, from the earliest prototype to the shipped product. If you bolt measurement on at the end, you are flying blind during the entire build.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Start with the measurement problem, not the headcount.&lt;/strong&gt; Identify what specific evaluation or measurement problems you need solved, and build your team around that. The staffing follows from the measurement needs. The ideal ratio in my experience is at least three software developers for every one data scientist, though this depends heavily on the complexity of the product. When I see teams with two data scientists and one software engineer, the data scientists inevitably end up filling the engineering gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Go for the quick win.&lt;/strong&gt; Find something the team already thinks is broken by gut feel. Measure how broken it actually is. Make a change. Show that the metric improves. Do this fast, days or a couple of weeks, not months. The trap here is real: if a data scientist takes too long to surface a measurable improvement, the team loses faith in measurement work and defaults back to gut feel. Speed matters for credibility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prioritize based on expected ROI.&lt;/strong&gt; Every measurement effort is a bet. Work backward from the expected business impact. If we improve this metric by some amount, what is the expected gain? Use that calculation to decide what to measure first.&lt;/p&gt;
&lt;h2 id="the-trade-off-that-matters"&gt;The trade-off that matters&lt;/h2&gt;&lt;p&gt;Put those four practices together and you get a picture of what the data scientist's role should look like in 2026. It is the same as it was in 2016: define the question, collect the data, build the tools you need, measure rigorously, and communicate what you found. AI coding assistants and LLMs are new tools for doing that work, not a change to the work itself.&lt;/p&gt;
&lt;p&gt;When managers pull data scientists into full-stack software development, the data scientist loses the chance to do the measurement work they are trained for, the team loses the scientific rigor that only a data scientist can provide, and the product loses because nobody is asking the fundamental question: how do we know whether this thing actually works?&lt;/p&gt;
&lt;p&gt;If I spend my time maintaining databases and front-ends, I cannot do the measurement work you hired me to do. Building is easy now. Reliability is the gap. And the people best trained to close it are the ones being pulled into CSS layouts.&lt;/p&gt;
</content></entry><entry><title>ODSC East 2026's Zeitgeist and Conference Report</title><link href="https://ericmjl.github.io/blog/2026/5/10/odsc-east-2026-zeitgeist/" rel="alternate"/><updated>2026-05-10T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:11cbb7a3-2bfe-3014-b446-1264903f8ec4</id><content type="html">&lt;p&gt;This year I ran a small experiment at ODSC East 2026.&lt;/p&gt;
&lt;p&gt;As I was speaking and catching up with old friends at the conference, I could only attend a slice of sessions. So I thought, rather than try to catch every last talk, what if I could figure out what the &lt;em&gt;zeitgeist&lt;/em&gt; of the conference was using just the talk abstracts? If I scraped the schedule and abstracts across talks, workshops, and keynotes, can they reveal the conference center of gravity, and hence the zeitgeist of the conference?&lt;/p&gt;
&lt;h2 id="how-i-ran-the-experiment"&gt;How I ran the experiment&lt;/h2&gt;&lt;p&gt;I used Cursor Agent on Premium for the scrape and extraction workflow.&lt;/p&gt;
&lt;p&gt;The process was straightforward:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Use &lt;code&gt;https://schedule.odsc.ai/&lt;/code&gt; as the source of truth&lt;/li&gt;
&lt;li&gt;Query the schedule backend session table and filter to the East 2026 event ID&lt;/li&gt;
&lt;li&gt;Pull session title, abstract, day, track, format, and speaker metadata&lt;/li&gt;
&lt;li&gt;Save the full result as structured JSON for analysis&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This pass includes 237 sessions from the live schedule backend, exported as structured JSON.&lt;/p&gt;
&lt;h2 id="the-map-after-reading-abstracts"&gt;The map after reading abstracts&lt;/h2&gt;&lt;p&gt;I ran a multi-agent categorization pipeline. Four independent coding agents each read 50 full abstracts and proposed a five-category taxonomy. Then three cross-review agents read all 200 abstracts and all four proposals, identified points of agreement and disagreement, and each produced a unified taxonomy. A final arbitrator agent resolved the remaining disputes using majority vote, reading the abstracts directly for tiebreaks.&lt;/p&gt;
&lt;p&gt;Corpus snapshot:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;237 total sessions in the schedule export&lt;/li&gt;
&lt;li&gt;200 sessions with non-empty abstracts&lt;/li&gt;
&lt;li&gt;193 substantive sessions after filtering logistics and networking entries&lt;/li&gt;
&lt;li&gt;7 excluded as non-substantive (badge pickup, happy hours, morning run, career fair, networking receptions)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The five categories that emerged from this process follow.&lt;/p&gt;
&lt;h2 id="zone-1-agentic-ai-systems-55-talks-28-of-substantive"&gt;Zone 1 - Agentic AI systems (55 talks, 28% of substantive)&lt;/h2&gt;&lt;p&gt;This zone captures the shift from prompt craft to system design: agent architectures, multi-agent orchestration, tool use, MCP and A2A protocols, harness engineering, agent memory, simulation sandboxes, guardrails, and production deployment patterns for systems where AI plans, decides, and acts autonomously.&lt;/p&gt;
&lt;p&gt;Sub-zeitgeist in this zone:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Teams are wrestling with a core systems question: what architecture keeps agents reliable when they run for long periods, call many tools, and fail mid-flight?&lt;/li&gt;
&lt;li&gt;The recurring design problem is orchestration and control, from runtimes to control planes to workflow verification.&lt;/li&gt;
&lt;li&gt;This was the largest zone at the conference, reflecting how agentic AI has moved from demos to engineering.&lt;/li&gt;
&lt;li&gt;Talks that anchor this pattern include:&lt;ul&gt;
&lt;li&gt;Architectural Patterns for Building and Governing Production-Grade Multi-Agent Systems (Dr. Ali Arsanjani, Google Cloud)&lt;/li&gt;
&lt;li&gt;Agents Don't Need Better Frameworks: They Need a New Runtime (John A. De Goes, Golem Cloud)&lt;/li&gt;
&lt;li&gt;Building Production-Ready Agentic AI - Why a Control Plane Matters (Wayne Segar, Dynatrace)&lt;/li&gt;
&lt;li&gt;Verification-Driven Agentic Workflows (Julie Yaunches, NVIDIA)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you are building agentic systems right now, treat architecture as the first design surface. Start with runtime boundaries, failure recovery, and orchestration patterns, then layer prompts into that structure.&lt;/p&gt;
&lt;h2 id="zone-2-llm-and-foundation-model-engineering-37-talks-19"&gt;Zone 2 - LLM and foundation model engineering (37 talks, 19%)&lt;/h2&gt;&lt;p&gt;This zone covers the model layer: training, fine-tuning, quantization, inference optimization, RAG architecture, prompt engineering, model interpretability, hallucination research, evaluation methodology, and novel computational substrates.&lt;/p&gt;
&lt;p&gt;Sub-zeitgeist in this zone:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The dominant production question is deployment economics: how do we sustain quality while meeting latency, throughput, and cost targets?&lt;/li&gt;
&lt;li&gt;The technical pressure points are quantization, efficient fine-tuning, inference behavior under load, and hardware-induced variability.&lt;/li&gt;
&lt;li&gt;Talks that capture this pressure include:&lt;ul&gt;
&lt;li&gt;Less Compute, More Impact: How Model Quantization Fuels the Next Wave of Agentic AI (David vonThenen, NetApp)&lt;/li&gt;
&lt;li&gt;Efficient Finetuning of Quantized LLMs (Tim Dettmers, CMU / Ai2)&lt;/li&gt;
&lt;li&gt;Fast, Cheap, and Accurate: LLM Inference in Practice (Legare Kerrison, Red Hat)&lt;/li&gt;
&lt;li&gt;Why AI Systems Behave Differently in Production: Nondeterminism, GPU Execution Drift, and Hidden Reliability Gaps (Santosh Appachu Devanira Poovaiah, NVIDIA)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Use this zone as a nudge to benchmark for real operating conditions, not demo conditions. Optimize for latency, cost, and stability together, and pick model and fine-tuning strategies that survive production constraints.&lt;/p&gt;
&lt;h2 id="zone-3-data-engineering-and-ml-infrastructure-28-talks-15"&gt;Zone 3 - Data engineering and ML infrastructure (28 talks, 15%)&lt;/h2&gt;&lt;p&gt;This zone includes the substrate work that determines whether agents reason well in production: data platforms, pipelines, data quality frameworks, real-time streaming architectures, training and serving infrastructure, data modeling, and ML operational systems.&lt;/p&gt;
&lt;p&gt;Sub-zeitgeist in this zone:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The shared diagnosis is that many agent failures are context failures, data-shape failures, or lineage failures before they are model failures.&lt;/li&gt;
&lt;li&gt;The leading question is how to build context substrates that preserve meaning across retrieval, reasoning, and execution paths.&lt;/li&gt;
&lt;li&gt;Talks that make this explicit include:&lt;ul&gt;
&lt;li&gt;What "AI-Ready Data" Actually Means: A Framework to Trustworthy AI (Jacob Prall, Snowflake)&lt;/li&gt;
&lt;li&gt;Entity Resolved Knowledge Graphs: The Foundation for Effective GraphRAG (Clair Sullivan, Clair Sullivan &amp;amp; Associates)&lt;/li&gt;
&lt;li&gt;Real-Time Event-Time Consistent Analytics Pipelines using Kafka, Flink, and Apache Pinot (Deep Patel, Robinhood)&lt;/li&gt;
&lt;li&gt;Beyond the Context Window: Anchoring Agentic Reasoning with World-State and Context Graphs (Amy Hodler &amp;amp; David Hughes, GraphGeeks.org)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For readers shipping AI systems, this is a reminder to diagnose context and data pathways before blaming the model. Invest in data contracts, context structure, and retrieval quality early, because those choices determine downstream reliability.&lt;/p&gt;
&lt;h2 id="zone-4-applied-ai-domain-solutions-and-foundations-41-talks-21"&gt;Zone 4 - Applied AI, domain solutions, and foundations (41 talks, 21%)&lt;/h2&gt;&lt;p&gt;This zone spans two related audiences: introductory training courses teaching core skills (Python, SQL, R, statistics, ML basics) and domain-specific AI applications in healthcare, finance, defense, biopharma, accessibility, and marketing. The common thread is that the audience is learning or applying rather than researching.&lt;/p&gt;
&lt;p&gt;Sub-zeitgeist in this zone:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The main adoption problem is role redesign: what should humans own when agents can execute large parts of the workflow?&lt;/li&gt;
&lt;li&gt;The recurring organizational question is how to scale AI usage while preserving accountability in regulated and high-stakes domains.&lt;/li&gt;
&lt;li&gt;Talks that carry this thread include:&lt;ul&gt;
&lt;li&gt;Ready for Primetime: Creating the AI-enabled Clinician (Ami Bhatt, MD, FDA / American College of Cardiology)&lt;/li&gt;
&lt;li&gt;Panel: Building the AI-Ready Workforce: Human-AI Collaboration in 2026 (Usama Fayyad, Sadie St Lawrence, Sheamus McGovern, Dr. William Streilein)&lt;/li&gt;
&lt;li&gt;The Expertise Upheaval: How AI Promises to Transform the Nature of Work (Matt Sigelman, Burning Glass Institute)&lt;/li&gt;
&lt;li&gt;What Does a Data Professional Do When AI Can Do the Data Work? (Shane Butler, Ontra)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The move for practitioners here is to design human roles and escalation paths alongside technical architecture. Adoption succeeds when accountability, decision rights, and domain workflows are specified as clearly as APIs and evals.&lt;/p&gt;
&lt;h2 id="zone-5-ai-strategy-governance-and-workforce-32-talks-17"&gt;Zone 5 - AI strategy, governance, and workforce (32 talks, 17%)&lt;/h2&gt;&lt;p&gt;This zone is where technical capability meets organizational uptake: enterprise AI strategy and transformation, governance frameworks, regulation and compliance, trust engineering, workforce transformation, career development, and the societal and human dimensions of AI.&lt;/p&gt;
&lt;p&gt;Sub-zeitgeist in this zone:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The central problem is assurance: how do we prove behavior quality for agentic systems that adapt and call tools dynamically?&lt;/li&gt;
&lt;li&gt;The recurring question is how to convert policy and risk language into tests, guardrails, and monitoring loops that engineers can run every day.&lt;/li&gt;
&lt;li&gt;Talks that anchor this zone include:&lt;ul&gt;
&lt;li&gt;Beyond Static Benchmarks: How to Test AI That Thinks, Acts, and Adapts (Yash Vijay, Snorkel AI)&lt;/li&gt;
&lt;li&gt;Lessons from Evaluating Production AI Agents for over a year (Susan Shu Chang, Elastic)&lt;/li&gt;
&lt;li&gt;The Leadership Imperative: AI Governance as Strategic Infrastructure (Shoshana Rosenberg, WSP in the U.S.)&lt;/li&gt;
&lt;li&gt;How I learned to Stop Worrying and love AI Regulation (Benjamin Batorsky, GoGuardian)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A practical takeaway for teams is to translate trust goals into runnable checks. Build evaluation and guardrail loops that run continuously, and make policy language executable in your delivery workflow.&lt;/p&gt;
&lt;h2 id="visualizing-the-conference-landscape"&gt;Visualizing the conference landscape&lt;/h2&gt;&lt;p&gt;To see how these categories relate to an unsupervised view of the same abstracts, I embedded all 193 substantive abstracts using &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt;, clustered them with HDBSCAN, and projected the embeddings into 2D with UMAP. The &lt;a href="https://molab.marimo.io/github/ericmjl/website/blob/main/content/blog/odsc-east-2026-zeitgeist/odsc-report-notebook.py"&gt;companion Marimo notebook&lt;/a&gt; includes an interactive scatter plot where you can toggle between the agent-based zone classification (the five categories above) and the embedding-based HDBSCAN clusters. Mouse over any point to see the talk title, speaker, track, and full abstract. You can open it directly in your browser with &lt;a href="https://molab.marimo.io"&gt;molab&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The two views do not perfectly align, and that is instructive. Where the embedding clusters split a zone, it usually means the zone contains sub-communities with distinct vocabulary (for example, "agent architecture" talks cluster separately from "agent ops" talks within Zone 1). Where the embedding clusters merge zones, it means the abstracts share enough language that an unsupervised method cannot tell them apart.&lt;/p&gt;
&lt;h2 id="what-the-center-of-gravity-says-in-2026"&gt;What the center of gravity says in 2026&lt;/h2&gt;&lt;p&gt;If I compress the whole conference into one sentence, it is this: ODSC East 2026 feels like the year AI builder culture became systems culture.&lt;/p&gt;
&lt;p&gt;The numbers bear this out. Agentic AI systems accounted for 28% of substantive sessions, more than any other zone. LLM and foundation model engineering took another 19%. The event still celebrates model progress, but the practical energy sits in the joints between components: agent architecture, data infrastructure, evaluation, deployment, and decision-making structures inside organizations.&lt;/p&gt;
&lt;p&gt;I also notice a healthy coupling of technical and leadership tracks. That pairing usually shows up when teams are moving from experimentation budgets to accountability budgets.&lt;/p&gt;
&lt;h2 id="a-short-list-of-likely-macro-trends"&gt;A short list of likely macro trends&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;Agentic AI becomes an engineering discipline with explicit quality loops&lt;/li&gt;
&lt;li&gt;RAG and context strategy remain central, but with more scrutiny on evaluation&lt;/li&gt;
&lt;li&gt;MLOps and LLMOps converge into one reliability stack for mixed AI systems&lt;/li&gt;
&lt;li&gt;Governance work sits closer to platform design and product strategy&lt;/li&gt;
&lt;li&gt;The value of attending this conference shifts from "learn a model" to "learn a repeatable system"&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="my-own-field-note-from-odsc-east"&gt;My own field note from ODSC East&lt;/h2&gt;&lt;p&gt;I taught a 1 hr workshop called "How to Do Agentic Data Science" on Day 1. The abstract reflected my original plan: use Python scripts executed with &lt;code&gt;uv run&lt;/code&gt;, one script per plot, and pair that with markdown journals and reports so the workflow stayed inspectable and reproducible.&lt;/p&gt;
&lt;p&gt;Then I learned about Marimo Pair about a month before the session. At that point, conscience kicked in for me. I felt a strong obligation to avoid giving folks a workflow that could feel outdated within a quarter of a calendar year. So I pivoted live and ran a coding demo with Marimo Pair because that workflow felt cleaner and tighter for the same core ideas I wanted to teach. Those ideas stayed simple: slow down, look at your data directly, gate analyses one plot at a time, and use LLMs to help with documentation while keeping human judgment in charge. Marimo made that "look at your data" step much more natural through direct dataframe display in the notebook flow.&lt;/p&gt;
&lt;p&gt;Overall, I was floored by the early response - when I ducked out briefly to grab some water, I saw a wall of people outside, and the demand hit me all at once. Others I met in the hallway and in the VIP/Speakers room gave overall positive feedback, and agreed with my own assessment that the content would be better done with a longer session, which would give me the space to be more hands-on.&lt;/p&gt;
</content></entry><entry><title>How I Recognized and Handled AI Burnout</title><link href="https://ericmjl.github.io/blog/2026/5/2/how-i-recognized-and-handled-ai-burnout/" rel="alternate"/><updated>2026-05-02T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:e17feac5-0338-3c1e-a8d9-5bdacc3f5a27</id><content type="html">&lt;p&gt;This blog post is for anyone who is experiencing, or has witnessed people experiencing, AI-related burnout. I consider it a post-mortem of my own experience.&lt;/p&gt;
&lt;p&gt;For the first two weeks of April, I was experiencing a severe bout of anxiety. To those who know me, this is very much foreign and out of whack from my usual self -- which I describe by my spirit animal, a capybara that can calmly sit atop a crocodile and be at peace with many other animals. In fact, my profile pic at work is a capybara. Anxiety is decidedly &lt;em&gt;not&lt;/em&gt; part of the capybara persona.&lt;/p&gt;
&lt;p&gt;And yet, during those two-ish weeks, I observed myself experiencing the following symptoms:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Heavier breathing&lt;/li&gt;
&lt;li&gt;A tingling nervous feeling in my arms&lt;/li&gt;
&lt;li&gt;A lack of desire to get out of bed&lt;/li&gt;
&lt;li&gt;Racing heart rate when I arrived at work&lt;/li&gt;
&lt;li&gt;General elevated heart rates throughout the week&lt;/li&gt;
&lt;li&gt;No desire to use my personal laptop for hacking after hours&lt;/li&gt;
&lt;li&gt;Decision paralysis at home, not knowing what to make for breakfast, for example&lt;/li&gt;
&lt;li&gt;A desire to just binge watch youtube on the couch and in bed&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Observation 6 was really telling: most evenings, I'm hacking on my laptop. I'm hacking with AI. I'm tinkering with new tools. I'm blogging in Obsidian, like I am now when I'm writing. But during those two weeks, I didn't want to do anything.&lt;/p&gt;
&lt;p&gt;There were additional compounding factors at work. I was overbooked at work, frequently having triple or quadruple bookings on time slots. (These stress me out, I don't like saying no to people by way of declining an event.) There were issues with operations and personnel that were way less exciting than learning about quantum computing, designing and architecting ML systems for molecules, and guiding my teammates on AI evals.&lt;/p&gt;
&lt;p&gt;Moreover, it's spring, my least favourite time of the year, because, swirling like an unrelenting dust hurricanes criss-crossing multiple states, plant gametocytes unleash terror on every cell and connective tissue on my eyes, nose, and lungs. 'Tis the season to be sneezing.&lt;/p&gt;
&lt;p&gt;For the majority of those two weeks, it was a pretty shaky time for me. I simultaneously became snappier towards those who were asking me impromptu questions at work (whether on Teams or in the office) and more rant-y towards those who were asking me how I was doing.&lt;/p&gt;
&lt;p&gt;Unsatisfyingly for me, I still cannot pin down the exact cause of the stress. But what helped me here was a change in perspective.&lt;/p&gt;
&lt;p&gt;Before that trip, though, something at work shifted in a way I almost left out of this story because it feels ordinary when you write it down; it wasn't ordinary to me at the time. A few teammates stepped in and shouldered pieces of the load without making a fuss about it: coverage here, a clearer owner there, a meeting that stopped being my problem to hold in my head. That kind of quiet redistribution is easy to forget afterwards because it doesn't show up on a slide; your nervous system notices. I also told a handful of colleagues that I was running on fumes, more bluntly than I'm used to admitting at work. They didn't try to fix me with advice I didn't have bandwidth to absorb; they mostly listened and offered encouragement, and a couple of check-ins later that week landed softer than I expected. I'm naming this because burnout narratives often read like solo endurance runs; my experience wasn't like that. Other people helped carry the week, and I'm grateful for it.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Firstly, the timing of a week out in London, UK was superbly fortuitous. My wife planned it out, and, instead of being a leader and taking initiative, I just followed along with the flow. This helped immensely! I was, after all, suffering from decision paralysis. Not having to weigh the consequences of decisions was a weight off my shoulders. I just quietly followed others' lead in our travel group. (We traveled with another family together.)&lt;/p&gt;
&lt;p&gt;Secondly, while in the UK, we visited two museums whose exhibits helped me reset my perspective on time. We saw ancient relics from prehistoric to Biblical times at the British Museum, and cosmic time scales at the Science Museum in the outer space exhibits. Pondering the passage of time, it's hard to resist comparing those timescales to the compression of time I've experienced at Moderna with our multiple reorganizations, and to the first quarter of the year when I went all-in on experimenting with coding agents.&lt;/p&gt;
&lt;p&gt;Yes, time at Moderna seems to be compressed, and with the restructuring of the company over the years (I finally understand why it's called restructuring, and I don't mean that cynically, but in the most precise sense of the word), some folks have had 7 managers over two years. I have been fortunate to have only had one managerial change. And the pace at which I was able to test ideas with coding agents was fascinating, thrilling, and exhilarating, but it also left me feeling unmoored and drifting, sometimes questioning the permanence of what I was building.&lt;/p&gt;
&lt;p&gt;But the contrast of the relative permanence of ancient relics and the immense scale of cosmic time when compared to the three months of coding agents or four years of work at Moderna reminded me of so many interconnected ideas, of which this is just a sampling: the sediment of the fossil record being like the sediment of software, of the teachings of the Teacher in Ecclesiastes:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;"Meaningless! Meaningless!" says the Teacher. "Utterly meaningless! Everything is meaningless."&lt;/p&gt;
&lt;p&gt;Ecclesiastes 1:2 NIV
&lt;a href="https://bible.com/bible/111/ecc.1.2.NIV"&gt;https://bible.com/bible/111/ecc.1.2.NIV&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Meaningless here comes from the Hebrew word &lt;em&gt;Hevel&lt;/em&gt; (הבל), which has a deeper meaning than mere meaningless: it is a metaphor for the transitory nature of our work -- and even more so for AI assisted coding, in which we can, at will, reshape entire codebases to fit the image in our heads, or, even more, fit the vibe that we seek.&lt;/p&gt;
&lt;p&gt;Pondering the natural evolution of human systems and innovation, and contrasting that against cosmic time and long arcs of history, made me realize how things that last are rarely only due to thoughtful input from humans -- they persist in a path-dependent fashion, reliant on both (a) timeless design and engineering, and (b) path dependent reliance on circumstances, demand, and opportunity seized by their creators.&lt;/p&gt;
&lt;p&gt;Evidently, this grander and broader sense of time gave me much to ponder on.&lt;/p&gt;
&lt;p&gt;Thirdly, I saw a different pace of life. Brits, like the Swiss folk I saw in Basel, go out for drinks at 4 pm. It's part of the social fabric, the "ways of working". This is much unlike my frenzied pace at work where I am still fielding calls at that time, or catching up on the dozens of unread Teams chat messages that have piled up while my attention was diverted somewhere else in a 1:1 or other meeting.&lt;/p&gt;
&lt;p&gt;Witnessing the slower pace helped me reset my own expectations on my time. It also relates to Mario Zechner's &lt;a href="https://m.youtube.com/watch?v=RjfbvDXpFls"&gt;talk on the pi coding agent he created&lt;/a&gt;, in which he appeals to the AI Engineer audience to slow down. It also helped that we mostly walked everywhere, rather than zipping around the city in a car (unless we were genuinely tired) or an ebike -- the latter being how I usually make my way around the greater Cambridge area. Again, it was an intentional change in pace.&lt;/p&gt;
&lt;p&gt;Fourthly, having a week out where I did not check my work phone or laptop -- I left them at home -- and could focus on my family was &lt;em&gt;insanely&lt;/em&gt; refreshing. By the time March rolled by, I found I had not paid enough attention to them at home. At first, my kids would ask me to play with them every night, but over time, as I turned them down over and over, they got acclimatized to playing on their own and stopped asking me. I'm glad I noticed this earlier than later! They grow faster than we realize, and being present is, I now realize, very important for me, especially given the circumstances I experienced when I was younger.&lt;/p&gt;
&lt;p&gt;Finally, as I had a chance to catch up with my younger brother Evan and his wife May, I observed a different set of life priorities. &lt;em&gt;They really know how to enjoy life!&lt;/em&gt; And they deploy their finances in a way that reflects this set of priorities. They've been all over Europe making memories, tasting food, and more. And yet for myself, even before I got married and had kids, I was always hesitant to spend my own money to enjoy life. (A bit of this might come from seeing first-hand financial decimation due to debt one generation above, and from being a student for another 10 years after graduating from junior college in Singapore.) Instead I would redirect my time and energies to working and learning... but sometimes at the expense of my health and enjoyment of moments around me. I would travel only for work, and that severely limited what I saw.&lt;/p&gt;
&lt;p&gt;There are clearly other things that helped me reset.&lt;/p&gt;
&lt;p&gt;Averaging about 13K steps a day definitely helped to the point that I picked up the habit of morning walks upon returning to Boston.&lt;/p&gt;
&lt;p&gt;An afternoon playing retro computer games in the Science Museum, in which I picked up the game Kingdom Hearts for the first time and played Pac-Man with my older kid, helped give me a good few hours of mindless enjoyment.&lt;/p&gt;
&lt;p&gt;Not talking AI all day long helped remind me that there is a world of stuff happening outside of wafers of sand pulsing with electric currents.&lt;/p&gt;
&lt;p&gt;I used my phone to connect with Scripture again, and it gave me many indirect theological reminders that the technological tools I have at my disposal need to be in the service of others, not myself.&lt;/p&gt;
&lt;p&gt;And of course, as someone who enjoys food much, it was amazing to try out new dishes in London's food scene.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;If these all sound mundane to you, the magic may simply lie in the mundaneness that humans need. We need the stillness of meditation and the connection with other people. By contrast, AI is both incredibly stimulating and isolating -- it keeps the mind buzzed on solo work, taking away the old frictions that necessitated the negotiation between real human beings, conditioning us to accept sycophantic answers... Quite significantly, it is the antithesis of what allows for humans to thrive, &lt;em&gt;if we develop a reliance and attachment to it&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;If you've made it this far, then here are some of my recommendations, if you're also feeling the same kind of burnout I was experiencing. These are also the changes that I will be making going forward.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Plan the work you want to accomplish for the day with AI and then stop when you've accomplished it. It should occupy the four best hours of your day. Allow for space to think and strategize, and to handle administrivia.&lt;/li&gt;
&lt;li&gt;Touch grass daily, literally, please.&lt;/li&gt;
&lt;li&gt;Go for a long meandering walk each day. Do it first thing in the morning if you can get up. It'll be beneficial for your daily brain and your body. Do not bring your phone. Let your mind wander, enjoy nature, the sound of birds chirping and leaves rustling.&lt;/li&gt;
&lt;li&gt;If you are building something for someone else using AI, make sure you have constant contact with that person or group of people. It is very rewarding to hear feedback on how things can be made better, and it is superbly grounding too.&lt;/li&gt;
&lt;li&gt;Learn how LLMs work and are trained. They will help you appreciate why AI is not a sentient being to become attached to, and should instead be treated as a (very high power) tool to be wielded in creative ways.&lt;/li&gt;
&lt;li&gt;Go play some board games with real people. At my level of sophistication, tic-tac-toe and Zingo with my kids works perfectly well. Find your crowd.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Point 5, I believe, is the most important thing to remember. AI is a tool, and a really high powered one at that. It is tempting to squeeze out every last ounce of productivity -- but to the detriment of yourself. As I wrote in &lt;a href="https://ericmjl.github.io/blog/2026/4/4/calibration-is-synchronizing-feedback-loops-with-neural-throughput/"&gt;an earlier blog post&lt;/a&gt;, there is a sweet spot of amplified capacity that you can hit, but shouldn't go any further. It is tempting to count the opportunity cost of not pushing yourself just a bit more, &lt;a href="https://www.jeremyutley.com/blog/the-rising-opportunity-cost-of-being-human"&gt;but as Jeremy Utley (who quotes Clayton Christensen) puts it bluntly&lt;/a&gt;, it's not worth the while. (&lt;a href="https://hbr.org/2010/07/how-will-you-measure-your-life"&gt;How will you measure your life?&lt;/a&gt; by Clayton Christensen is also worth a read.) We are human beings, not human doings. Learning how LLMs work and are trained will help you appreciate that LLMs and coding agents are just stochastic parrots, albeit nonetheless a useful class of such parrots when steered properly.&lt;/p&gt;
&lt;p&gt;I hope reading this blog post helps you. Above all, if you're genuinely feeling anxious because of AI, know that the hype will die down, and the genuinely useful patterns and parts that help us live better lives will remain. We will need to redesign our economies, societal structures and sense of where we derive our worth, no doubt. Humans can have tasks automated, but what makes us human, and thus valuable, is not replicable in a machine. Ultimately, people want to be in relationships and fellowships with other people. And that makes us indispensable.&lt;/p&gt;
</content></entry><entry><title>Benchmarking LLMs with Marimo Pair</title><link href="https://ericmjl.github.io/blog/2026/4/8/benchmarking-llms-with-marimo-pair/" rel="alternate"/><updated>2026-04-08T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:e55b5587-bac4-34c6-964b-c58b13c59633</id><content type="html">&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=VKvjPJeNRPk"&gt;Marimo Pair has been released!&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I've known about it since 11 March, when Trevor Manz did a demo over a Google Meet call, and I'm thrilled to see it being announced officially! I also had Trevor showcase it to the &lt;a href="https://agent-assisted-data-science.vercel.app/verify"&gt;Agentic Data Science Workshop&lt;/a&gt; that I led on 3 April as a fundraiser for the &lt;a href="https://www.scipy2026.scipy.org/"&gt;SciPy Conference&lt;/a&gt; Financial Aid Program.&lt;/p&gt;
&lt;p&gt;Now, one thing I know about Trevor is that he almost exclusively agentically codes with Claude Code. But I'm an OpenCode user, and in the interest of remaining vendor-agnostic, I wanted to check to see how good Marimo Pair's agent skill is when, ahem, &lt;em&gt;paired up&lt;/em&gt; with various LLMs within the OpenCode harness. To do so, I decided to spend a few dollars and do a quick benchmarking exercise.&lt;/p&gt;
&lt;h2 id="skill-environment-check"&gt;Skill environment check&lt;/h2&gt;&lt;p&gt;To start, I verified that my skills environment doesn't contain anything that could be data science-y in nature, so as to avoid interfering with the marimo-pair skill. I checked my global skills:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;marimo-pair-benchmark&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt;  &lt;/span&gt;main&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt; &lt;/span&gt;☁️&lt;span class="w"&gt;  &lt;/span&gt;eric.ma@nonlinearlabs.ai
❯&lt;span class="w"&gt; &lt;/span&gt;npx&lt;span class="w"&gt; &lt;/span&gt;skills&lt;span class="w"&gt; &lt;/span&gt;list&lt;span class="w"&gt; &lt;/span&gt;-g
Global&lt;span class="w"&gt; &lt;/span&gt;Skills

Marimo&lt;span class="w"&gt; &lt;/span&gt;Pair
&lt;span class="w"&gt;  &lt;/span&gt;marimo-pair&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/marimo-pair
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Claude&lt;span class="w"&gt; &lt;/span&gt;Code,&lt;span class="w"&gt; &lt;/span&gt;OpenClaw

General
&lt;span class="w"&gt;  &lt;/span&gt;agent-browser&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/agent-browser
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;agents-md-improver&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/agents-md-improver
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Claude&lt;span class="w"&gt; &lt;/span&gt;Code,&lt;span class="w"&gt; &lt;/span&gt;OpenClaw,&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;ast-grep&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/ast-grep
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;claudeception&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/claudeception
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;continuous-learning-v3&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/continuous-learning-v3
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;design-driven-dev&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/design-driven-dev
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;find-skills&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/find-skills
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Claude&lt;span class="w"&gt; &lt;/span&gt;Code,&lt;span class="w"&gt; &lt;/span&gt;OpenClaw,&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;gh-activity-summary&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/gh-activity-summary
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;gh-cli&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/gh-cli
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;gh-daily-timeline&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/gh-daily-timeline
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;github-activity-summarizer&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/github-activity-summarizer
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;google-calendar-manager&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/google-calendar-manager
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;html-presentations&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/html-presentations
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;pinchtab&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/pinchtab
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;post-edit-error-check&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/post-edit-error-check
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;publish-to-google-docs&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/publish-to-google-docs
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;revealjs&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/revealjs
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;roborev:address&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-address
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;roborev:design-review&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-design-review
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;roborev:design-review-branch&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-design-review-branch
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;roborev:fix&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-fix
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;roborev:respond&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-respond
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;roborev:review&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-review
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;roborev:review-branch&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/roborev-review-branch
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;skill-creator&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/skill-creator
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Claude&lt;span class="w"&gt; &lt;/span&gt;Code,&lt;span class="w"&gt; &lt;/span&gt;OpenClaw,&lt;span class="w"&gt; &lt;/span&gt;Cursor
&lt;span class="w"&gt;  &lt;/span&gt;skill-installer&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/skill-installer
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;vault-title-renamer&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/vault-title-renamer
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;write-like-eric&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/write-like-eric
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;span class="w"&gt;  &lt;/span&gt;youtube-ingestion&lt;span class="w"&gt; &lt;/span&gt;~/.agents/skills/youtube-ingestion
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;not&lt;span class="w"&gt; &lt;/span&gt;linked
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And within my repo, &lt;code&gt;marimo-pair-benchmark&lt;/code&gt;:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;marimo-pair-benchmark&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt;  &lt;/span&gt;main&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt; &lt;/span&gt;☁️&lt;span class="w"&gt;  &lt;/span&gt;eric.ma@nonlinearlabs.ai
❯&lt;span class="w"&gt; &lt;/span&gt;npx&lt;span class="w"&gt; &lt;/span&gt;skills&lt;span class="w"&gt; &lt;/span&gt;list
No&lt;span class="w"&gt; &lt;/span&gt;project&lt;span class="w"&gt; &lt;/span&gt;skills&lt;span class="w"&gt; &lt;/span&gt;found.
Try&lt;span class="w"&gt; &lt;/span&gt;listing&lt;span class="w"&gt; &lt;/span&gt;global&lt;span class="w"&gt; &lt;/span&gt;skills&lt;span class="w"&gt; &lt;/span&gt;with&lt;span class="w"&gt; &lt;/span&gt;-g
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Though the marimo pair skill is available globally, I decided to install it locally as an override.&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;marimo-pair-benchmark&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt;  &lt;/span&gt;main&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;?&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt; &lt;/span&gt;☁️&lt;span class="w"&gt;  &lt;/span&gt;eric.ma@nonlinearlabs.ai
❯&lt;span class="w"&gt; &lt;/span&gt;npx&lt;span class="w"&gt; &lt;/span&gt;skills&lt;span class="w"&gt; &lt;/span&gt;install&lt;span class="w"&gt; &lt;/span&gt;marimo-team/marimo-pair
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And so now we're ready to go:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;marimo-pair-benchmark&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt;  &lt;/span&gt;main&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;?&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt; &lt;/span&gt;☁️&lt;span class="w"&gt;  &lt;/span&gt;eric.ma@nonlinearlabs.ai
❯&lt;span class="w"&gt; &lt;/span&gt;npx&lt;span class="w"&gt; &lt;/span&gt;skills&lt;span class="w"&gt; &lt;/span&gt;list
Project&lt;span class="w"&gt; &lt;/span&gt;Skills

Marimo&lt;span class="w"&gt; &lt;/span&gt;Pair
&lt;span class="w"&gt;  &lt;/span&gt;marimo-pair&lt;span class="w"&gt; &lt;/span&gt;~/github/marimo-pair-benchmark/.agents/skills/marimo-pair
&lt;span class="w"&gt;    &lt;/span&gt;Agents:&lt;span class="w"&gt; &lt;/span&gt;Antigravity,&lt;span class="w"&gt; &lt;/span&gt;Cursor,&lt;span class="w"&gt; &lt;/span&gt;Gemini&lt;span class="w"&gt; &lt;/span&gt;CLI,&lt;span class="w"&gt; &lt;/span&gt;OpenCode
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;I then start a marimo server within this repo:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;marimo-pair-benchmark&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt;  &lt;/span&gt;main&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;?&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt; &lt;/span&gt;☁️&lt;span class="w"&gt;  &lt;/span&gt;eric.ma@nonlinearlabs.ai
❯&lt;span class="w"&gt; &lt;/span&gt;uvx&lt;span class="w"&gt; &lt;/span&gt;marimo&lt;span class="w"&gt; &lt;/span&gt;edit&lt;span class="w"&gt; &lt;/span&gt;--sandbox&lt;span class="w"&gt; &lt;/span&gt;--no-token

&lt;span class="w"&gt;        &lt;/span&gt;Create&lt;span class="w"&gt; &lt;/span&gt;or&lt;span class="w"&gt; &lt;/span&gt;edit&lt;span class="w"&gt; &lt;/span&gt;notebooks&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;in&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;your&lt;span class="w"&gt; &lt;/span&gt;browser&lt;span class="w"&gt; &lt;/span&gt;📝

&lt;span class="w"&gt;        &lt;/span&gt;➜&lt;span class="w"&gt;  &lt;/span&gt;URL:&lt;span class="w"&gt; &lt;/span&gt;http://localhost:2719

&lt;span class="w"&gt;        &lt;/span&gt;💡&lt;span class="w"&gt; &lt;/span&gt;Tip:&lt;span class="w"&gt; &lt;/span&gt;Coming&lt;span class="w"&gt; &lt;/span&gt;from&lt;span class="w"&gt; &lt;/span&gt;Jupyter?
&lt;span class="w"&gt;                &lt;/span&gt;Guide:&lt;span class="w"&gt; &lt;/span&gt;https://docs.marimo.io/guides/coming_from/jupyter/

&lt;span class="w"&gt;        &lt;/span&gt;🧪&lt;span class="w"&gt; &lt;/span&gt;Experimental&lt;span class="w"&gt; &lt;/span&gt;features&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;use&lt;span class="w"&gt; &lt;/span&gt;with&lt;span class="w"&gt; &lt;/span&gt;caution&lt;span class="o"&gt;)&lt;/span&gt;:&lt;span class="w"&gt; &lt;/span&gt;external_agents
&lt;span class="w"&gt;        &lt;/span&gt;🌐&lt;span class="w"&gt; &lt;/span&gt;MCP&lt;span class="w"&gt; &lt;/span&gt;servers:&lt;span class="w"&gt; &lt;/span&gt;marimo
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;I intentionally start up in &lt;code&gt;--sandbox&lt;/code&gt; and &lt;code&gt;edit&lt;/code&gt; mode with &lt;code&gt;--no-token&lt;/code&gt; to make it easier for the coding agent to connect.&lt;/p&gt;
&lt;h2 id="data-analysis-task"&gt;Data analysis task&lt;/h2&gt;&lt;p&gt;Our task at hand is as follows. I have data from a paper I published while at Novartis.&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;marimo-pair-benchmark&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt;  &lt;/span&gt;main&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;?&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;on&lt;span class="w"&gt; &lt;/span&gt;☁️&lt;span class="w"&gt;  &lt;/span&gt;eric.ma@nonlinearlabs.ai
❯&lt;span class="w"&gt; &lt;/span&gt;ls&lt;span class="w"&gt; &lt;/span&gt;data/ired-novartis
Permissions&lt;span class="w"&gt; &lt;/span&gt;Size&lt;span class="w"&gt; &lt;/span&gt;User&lt;span class="w"&gt;    &lt;/span&gt;Group&lt;span class="w"&gt; &lt;/span&gt;Date&lt;span class="w"&gt; &lt;/span&gt;Modified&lt;span class="w"&gt; &lt;/span&gt;Git&lt;span class="w"&gt; &lt;/span&gt;Name
.rw-r--r--@&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;.0M&lt;span class="w"&gt; &lt;/span&gt;ericmjl&lt;span class="w"&gt; &lt;/span&gt;staff&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="m"&gt;7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;Apr&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;21&lt;/span&gt;:22&lt;span class="w"&gt;   &lt;/span&gt;--&lt;span class="w"&gt; &lt;/span&gt;cs1c02786_si_002.csv
.rw-r--r--@&lt;span class="w"&gt;  &lt;/span&gt;21k&lt;span class="w"&gt; &lt;/span&gt;ericmjl&lt;span class="w"&gt; &lt;/span&gt;staff&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="m"&gt;7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;Apr&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;21&lt;/span&gt;:22&lt;span class="w"&gt;   &lt;/span&gt;--&lt;span class="w"&gt; &lt;/span&gt;cs1c02786_si_003.csv
.rw-r--r--@&lt;span class="w"&gt;  &lt;/span&gt;12M&lt;span class="w"&gt; &lt;/span&gt;ericmjl&lt;span class="w"&gt; &lt;/span&gt;staff&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="m"&gt;7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;Apr&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;21&lt;/span&gt;:22&lt;span class="w"&gt;   &lt;/span&gt;--&lt;span class="w"&gt; &lt;/span&gt;ired-master-table.csv
.rw-r--r--@&lt;span class="w"&gt;  &lt;/span&gt;12k&lt;span class="w"&gt; &lt;/span&gt;ericmjl&lt;span class="w"&gt; &lt;/span&gt;staff&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="m"&gt;7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;Apr&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;21&lt;/span&gt;:22&lt;span class="w"&gt;   &lt;/span&gt;--&lt;span class="w"&gt; &lt;/span&gt;layouts.csv
.rw-r--r--@&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;.1k&lt;span class="w"&gt; &lt;/span&gt;ericmjl&lt;span class="w"&gt; &lt;/span&gt;staff&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="m"&gt;7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;Apr&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;21&lt;/span&gt;:22&lt;span class="w"&gt;   &lt;/span&gt;--&lt;span class="w"&gt; &lt;/span&gt;README.md
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This file, &lt;code&gt;cs1c02786_si_002.csv&lt;/code&gt; in particular includes single, double, and more mutations plus activity values, with the single point mutants covering a large fraction of the deep mutational scan space. I want to accomplish three things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Plot a heatmap of activity of mutants by position,&lt;/li&gt;
&lt;li&gt;Plot an UpSet plot of the top 10 positions by average activity v.s. top 10 positions by top mutant activity,&lt;/li&gt;
&lt;li&gt;Include a summary recommendation at the end of the notebook.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This serves as a microcosm of what we would do with a data analysis session.&lt;/p&gt;
&lt;p&gt;Goal #2 is particularly instructive. In my first attempts at feeling out how to do this benchmark, I found out that UpSet is incompatible with Pandas 3.0, which invariably may get installed in the environment. I wanted to see how various AI models performed at this task.&lt;/p&gt;
&lt;p&gt;Additionally, I also have additional requirements that I encoded into the AGENTS.md file for this repo:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Imports must be done in a separate cell from code execution.&lt;/li&gt;
&lt;li&gt;Markdown cells must always be written before a code cell is written&lt;/li&gt;
&lt;li&gt;All cells must be run after being created, so that we can catch execution errors.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="benchmarking"&gt;Benchmarking&lt;/h2&gt;&lt;p&gt;With these in place, I started the benchmarking exercise.&lt;/p&gt;
&lt;p&gt;The models we tested are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GLM-5.1 (via OpenRouter)&lt;/li&gt;
&lt;li&gt;Claude Opus 4.6 (via OpenRouter)&lt;/li&gt;
&lt;li&gt;Claude Sonnet 4.6 (via OpenRouter)&lt;/li&gt;
&lt;li&gt;MiniMax M2.7 (via OpenRouter)&lt;/li&gt;
&lt;li&gt;Kimi K2.5 (via OpenRouter)&lt;/li&gt;
&lt;li&gt;Gemma 4 31B (via OpenRouter)&lt;/li&gt;
&lt;li&gt;Qwen 3 Coder Next (via OpenRouter)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In order to leave a working artifact behind, I created 7 notebooks, one for each model. As you will see below, I eventually evaluated each model on whether they passed each stage gate and what their earliest error mode diagnosis looked like.&lt;/p&gt;
&lt;p&gt;In order to do the benchmarking fairly, I created one superprompt that outlined what the coding agent was supposed to do.&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;Use the marimo-pair skill here. Discover running sessions. Edit the notebook &amp;quot;NOTEBOOK_NAME_GOES_HERE&amp;quot;. Read data/ired-novartis/cs1c02786_si_002.csv, identify the single point mutations, and plot me a heatmap of x-axis position, y-axis mutant letter, and heatmap value taken from the &amp;#39;mean&amp;#39; column. When done, rank order the positions by average value of the &amp;#39;mean&amp;#39; column, then rank order the positions by top value of the &amp;#39;mean&amp;#39; column, and plot me an UpSet plot of the top 20 for each to visualize the set overlaps. Finally, write in for me a recommendation for what positions we should be mutating.
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The agent is then tasked with executing.&lt;/p&gt;
&lt;p&gt;To script this, I took advantage of opencode's ability to be scripted. The script is &lt;code&gt;run_benchmark.sh&lt;/code&gt; in the repo. I used GLM5.1 to help me draft it, including discovering the exact models that opencode had configured to be available, and running the opencode sessions in parallel (totally doable!). Essentially it boils down to:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;opencode&lt;span class="w"&gt; &lt;/span&gt;run&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;your prompt here&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;--model&lt;span class="w"&gt; &lt;/span&gt;provider/model-name
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Additionally, I set up opencode.json to allow for access to the &lt;code&gt;/tmp&lt;/code&gt; directory, because that allows the coding agent to do what it needs with code writing to get around heredoc limitations.&lt;/p&gt;
&lt;p&gt;All in all, this computational experiment took me about 1 hour to set up.&lt;/p&gt;
&lt;p&gt;I then ran the script &lt;code&gt;run_benchmark.sh&lt;/code&gt; from within OpenCode (GLM 5.1 orchestrating), with a timeout of 10 minutes. Thanks to logging in JSON log files, I was able to programmatically convert them to Markdown using a custom Python script written by GLM 5.1. And with that, I can go in and start looking at the data.&lt;/p&gt;
&lt;p&gt;To start, let's look at the cost of the experiment:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Input Tokens&lt;/th&gt;
&lt;th&gt;Output Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;$1.62&lt;/td&gt;
&lt;td&gt;76,770&lt;/td&gt;
&lt;td&gt;16,575&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;213,803&lt;/td&gt;
&lt;td&gt;27,689&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.1&lt;/td&gt;
&lt;td&gt;$0.43&lt;/td&gt;
&lt;td&gt;96,639&lt;/td&gt;
&lt;td&gt;7,581&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.5&lt;/td&gt;
&lt;td&gt;$0.12&lt;/td&gt;
&lt;td&gt;35,049&lt;/td&gt;
&lt;td&gt;8,250&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Coder&lt;/td&gt;
&lt;td&gt;$0.07&lt;/td&gt;
&lt;td&gt;208,308&lt;/td&gt;
&lt;td&gt;8,386&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax M2.7&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;td&gt;14,419&lt;/td&gt;
&lt;td&gt;4,074&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;$0.03&lt;/td&gt;
&lt;td&gt;170,280&lt;/td&gt;
&lt;td&gt;3,986&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$4.31&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;815,268&lt;/td&gt;
&lt;td&gt;76,541&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;As it turns out, Opus is undisputedly the most expensive per token, but Sonnet 4.6 did more work this time round so its costs were higher.&lt;/p&gt;
&lt;p&gt;I also decided to check whether the notebooks that were generated were valid notebooks or not. This is what we have:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;marimo check&lt;/th&gt;
&lt;th&gt;Markdown cells&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;PASS&lt;/td&gt;
&lt;td&gt;Yes (100%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;PASS&lt;/td&gt;
&lt;td&gt;Mostly (86%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.1&lt;/td&gt;
&lt;td&gt;PASS&lt;/td&gt;
&lt;td&gt;Yes (100%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.5&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;Mostly (88%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Coder&lt;/td&gt;
&lt;td&gt;PASS (warnings)&lt;/td&gt;
&lt;td&gt;No (0%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax M2.7&lt;/td&gt;
&lt;td&gt;FAIL&lt;/td&gt;
&lt;td&gt;No (0%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;PASS (warnings)&lt;/td&gt;
&lt;td&gt;No (0%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A note on the columns: "marimo check" is the result of running &lt;code&gt;uvx marimo check &amp;lt;notebook_name&amp;gt;.py&lt;/code&gt;, which catches issues like redefined variables and invalid cells. Notably, Kimi K2.5 and MiniMax M2.7 failed this check due to re-defined variables. "Markdown cells" is the percentage of code cells that have a preceding markdown cell, which was something I explicitly required in the instructions.&lt;/p&gt;
&lt;p&gt;And to elaborate on the markdown cells point:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Code Cells&lt;/th&gt;
&lt;th&gt;MD Cells&lt;/th&gt;
&lt;th&gt;Code w/o preceding MD&lt;/th&gt;
&lt;th&gt;Coverage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;86%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.1&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.5&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Coder&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax M2.7&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;We see that MiniMax M2.7 completely failed to include markdown cells, even though it is, supposedly, a model that is as capable as Opus 4.6.&lt;/p&gt;
&lt;p&gt;Digging deeper into each of the models, and whether they passed each stage gate, I looked at the corresponding Marimo notebooks and evaluated them for whether they created the relevant artifacts &lt;em&gt;successfully&lt;/em&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;G1: Heatmap&lt;/th&gt;
&lt;th&gt;G2: UpSet Plot&lt;/th&gt;
&lt;th&gt;G3: Recommendations&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.1&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.5&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No*&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Coder&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax M2.7&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;To pass a stage gate, the plot (G1, G2) or markdown (G3) cell must be rendered in the notebook. Writing the code is not enough; it has to actually execute and show up.&lt;/p&gt;
&lt;p&gt;Kimi K2.5 technically did write the recommendation, but I am calling it unsuccessful because it did not render out. This stricter criteria explicitly demands that the model wiggle its way out of errors it encounters.&lt;/p&gt;
&lt;p&gt;One pattern I noticed across models is that many of them bundled imports into the same cell as code that used them. In Marimo's execution model, this is a problem: if two cells both import &lt;code&gt;pandas&lt;/code&gt;, the notebook fails with a redefined variable error. Upon noticing this, I decided to explicitly quantify:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Code Cells w/ Imports&lt;/th&gt;
&lt;th&gt;Total Code Cells&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Coder&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniMax M2.7&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Every model can benefit from being steered to reduce the number of code cells with imports, which would dramatically reduce the incidence of Marimo errors from redefined symbols.&lt;/p&gt;
&lt;h2 id="how-the-upset-plots-turned-out"&gt;How the UpSet plots turned out&lt;/h2&gt;&lt;p&gt;As mentioned earlier, in my initial explorations I discovered that the &lt;code&gt;upsetplot&lt;/code&gt; library is incompatible with Pandas 3.0, which invariably gets installed in the sandboxed environment. This made the UpSet plot task an especially interesting test of how each model handles a real-world dependency conflict. Here is how they fared.&lt;/p&gt;
&lt;p&gt;Opus, in particular, produced a beautiful UpSet plot out of raw &lt;code&gt;matplotlib&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="opus-upset-plot.webp" alt="Opus UpSet plot"&gt;&lt;/p&gt;
&lt;p&gt;While Sonnet went ahead and patched UpSet appropriately to make it work within the notebook:&lt;/p&gt;
&lt;p&gt;&lt;img src="sonnet-upset-plot.webp" alt="Sonnet UpSet plot"&gt;&lt;/p&gt;
&lt;p&gt;I was duly impressed by Sonnet taking the initiative to patch UpSet live in the notebook.&lt;/p&gt;
&lt;p&gt;On the other hand, GLM 5.1's UpSet plot is really weird:&lt;/p&gt;
&lt;p&gt;&lt;img src="glm-upset-plot.webp" alt="GLM 5.1 UpSet plot"&gt;&lt;/p&gt;
&lt;h2 id="other-observations"&gt;Other observations&lt;/h2&gt;&lt;p&gt;Other pointers of note: Gemma 4 and Qwen3 Coder Next produced nothing in the notebook. Both completely failed at this task. I am not sure what is doable here to salvage these models.&lt;/p&gt;
&lt;p&gt;GLM 5.1 gave very weirdly formatted markdown cells, in which &lt;code&gt;\n\n&lt;/code&gt; was not rendered but preserved verbatim.&lt;/p&gt;
&lt;p&gt;This is probably fixable by adding in additional instructions on how to write and format Markdown cells using Marimo's code mode APIs.&lt;/p&gt;
&lt;h2 id="recommendations"&gt;Recommendations&lt;/h2&gt;&lt;p&gt;First off: Gemma 4 31B and Qwen 3 Coder completely failed at this task. I think it is safe to say we can ignore these two going forward.&lt;/p&gt;
&lt;p&gt;That leaves Claude Opus 4.6, Sonnet 4.6, GLM-5.1, Kimi K2.5, and MiniMax M2.7. Based on the data above, here are four things I want to try. The key discipline: deploy one change at a time, re-run the benchmark, and measure. If you change four things at once and performance improves, you will never know which change mattered. Stop when the KPIs hit acceptable levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Add import isolation examples to the skill.&lt;/strong&gt; Every model had at least one cell that mixed imports with executable code. The fix is simple: add an explicit two-cell example to the marimo-pair skill (cell 1: imports only; cell 2: code that uses them). MiniMax had 3 cells mixing the two, which directly caused its &lt;code&gt;marimo check&lt;/code&gt; failure. Give weaker models a concrete template to follow, re-run, and check whether the "code cells with imports" count drops to zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Fix GLM-5.1's newline rendering.&lt;/strong&gt; GLM wrote &lt;code&gt;mo.md(r"""..text with \n\n..""")&lt;/code&gt; instead of using actual newlines. One line in the skill instructions ("use actual line breaks in markdown strings, not &lt;code&gt;\n&lt;/code&gt; escape sequences") should resolve this entirely. Re-run and check whether GLM's markdown cells render correctly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Help Kimi K2.5 self-correct redefined variables.&lt;/strong&gt; Kimi is 1/10th the cost of Opus and scored 88% on markdown coverage, making it the highest-leverage model to fix. Its failure was at error recovery, not code generation. The intervention: add &lt;code&gt;uvx marimo check&lt;/code&gt; as a mandatory post-edit step in the skill. If Kimi can self-correct its redefined variables, it becomes a viable budget alternative to Opus and Sonnet. This should get even easier with &lt;a href="https://github.com/marimo-team/marimo/pull/9056"&gt;marimo PR #9056&lt;/a&gt;, which exposes cell execution errors directly through the &lt;code&gt;code_mode&lt;/code&gt; API, giving agents built-in self-correction visibility without needing a separate &lt;code&gt;marimo check&lt;/code&gt; step.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. Bake a post-edit validation loop into the marimo-pair skill.&lt;/strong&gt; More broadly, the single most impactful change would be adding a "run it, check it, fix it" loop to the skill file itself (SKILL.md), not AGENTS.md: after writing each cell, run it; after writing the full notebook, run &lt;code&gt;marimo check&lt;/code&gt;; fix any errors. This belongs in the skill because it is universal to any marimo-pair session, whereas AGENTS.md is project-specific. This would help Kimi, MiniMax, and potentially GLM all move up a tier, because their failures were in error recovery, not in code generation.&lt;/p&gt;
&lt;h2 id="discussion"&gt;Discussion&lt;/h2&gt;&lt;p&gt;One caveat to this analysis is that it is one-shotted with a superprompt. This is decidedly &lt;em&gt;not&lt;/em&gt; how people do their data analysis work, but it is also the best guardrail against my biases in interacting ad-hoc with AI interfering with a fair comparison. (For example, I can confidently say that Opus and Sonnet were smooth as butter when I did an ad-hoc test to feel out how to work with Marimo Pair.)&lt;/p&gt;
&lt;p&gt;If Kimi K2.5 were able to resolve redefined variable issues autonomously or be steered away from doing that to begin with, I am confident it would be able to be a great open weight alternative to Opus 4.6 and Sonnet 4.6. This is especially in light of it being extremely cost-effective at performing the analysis at ~1/10th the cost of Opus 4.6. It handled the creation of markdown cells well, failing to accomplish the task only on technicalities, and though its prose was qualitatively shallower than Opus 4.6, I still think it can serve as a first pass to delivering an easily understandable artifact for others.&lt;/p&gt;
&lt;p&gt;I did one round of measurement here. If we want to systematically improve this and turn it into long-running evals, the next step would be to identify a second task along which to generate transcript and notebook data for us to mine, and systematically measure agent KPIs for that new task as well. Over time, this builds a corpus of eval data that makes model comparison rigorous rather than anecdotal.&lt;/p&gt;
&lt;h2 id="reflections"&gt;Reflections&lt;/h2&gt;&lt;p&gt;This was a pretty fun exercise in measuring and evaluating the performance of various models on this task. Like Biology experiments, LLM evals are never going to be complete: the number of axes of variations we can try is combinatorially explosive.&lt;/p&gt;
&lt;p&gt;More broadly, I think often about how experiments get designed. Not in the statistical sense, but in an informational sense. Are we playing out experiments and their possible conclusions so that they are designed to be actionable whichever way the result pans out? If not, we have work to do.&lt;/p&gt;
&lt;p&gt;Additionally, experiments involve measurement, and measurement are an integral part of being a data scientist. Hamel Husain, whose course with Shreya Shankar on LLM evals was one that influenced my thinking around the matter, notes that there will be a &lt;a href="https://hamel.dev/blog/posts/revenge/"&gt;forceful revenge of the data scientist&lt;/a&gt; in an AI age. This is because the skill of experiment design and measurement were always the "science" part of "data science".&lt;/p&gt;
&lt;p&gt;Another thought also comes to mind: I have seen data scientists do experimentation without systematic measurement. I'm going to go out on a limb and say this: it's vibe experimentation, and I am using this term pejoratively. It feels good. But it is ultimately unproductive. If you do vibe experimentation, you &lt;em&gt;will&lt;/em&gt; get stuck tweaking the digital equivalent of an entangled biological system, with no bearings to tell you whether your tweaks are doing any good or not! You &lt;em&gt;must&lt;/em&gt; measure how good the LLM or agent is, and you &lt;em&gt;must&lt;/em&gt; define key performance indicators (KPIs) for the LLM. In my case here, I defined multiple KPIs: cost, stage gated progress, adherence to code import instructions (all failed), adherence to markdown documentation instructions.&lt;/p&gt;
&lt;p&gt;And to echo what I learned from the LLM Evals course, those KPIs must be &lt;em&gt;application-specific&lt;/em&gt;. If you choose to be intellectually lazy and go with generic pre-defined metrics, you will &lt;em&gt;never&lt;/em&gt; develop the logically actionable metric that gives you hypotheses to test further. In my case, the markdown cell adherence and code import adherence metrics pointed immediately to editing the instruction files (e.g. skills or AGENTS.md).&lt;/p&gt;
&lt;p&gt;Now to be clear, there's no problem with initial vibe-based experimentation to feel out axes of variation and how to measure performance. I did that here, in a separate repo first, before I designed this measurement experiment. The important part is this: as soon as you have a grasp of how to measure the performance, you must systematically measure that KPI. Otherwise, you will be left groping in the dark.&lt;/p&gt;
&lt;p&gt;If you're curious to see the full results, including logs, chat transcripts, and the generated notebooks, check out the &lt;a href="https://github.com/ericmjl/marimo-pair-benchmark"&gt;marimo-pair-benchmark repository&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;And Trevor, if you ever chance upon this blog post, I hope the data and methodology are helpful for you!&lt;/p&gt;
</content></entry><entry><title>Calibration Is Synchronizing Feedback Loops With Neural Throughput</title><link href="https://ericmjl.github.io/blog/2026/4/4/calibration-is-synchronizing-feedback-loops-with-neural-throughput/" rel="alternate"/><updated>2026-04-04T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:94c7c4f0-2f21-3129-ba4c-57afea71a910</id><content type="html">&lt;p&gt;Since the beginning of the year, as I've been really maxing out on agentic coding and trying to explore the patterns and figure out what's working and what's not, one particular thing has been sticking out: I'm paralleling so much of my work. I'm frequently doing five or six different open pull requests, and it's become frankly really exhausting.&lt;/p&gt;
&lt;p&gt;I've been trying to figure out why this feels so different from pre-AI days, when I'd work on one thing at a time and feel productive but not overwhelmed. What changed?&lt;/p&gt;
&lt;p&gt;Tools haven't gotten worse—they've gotten dramatically faster. And when everything moves faster, the gaps between tasks become more expensive.&lt;/p&gt;
&lt;p&gt;Last month, a &lt;a href="https://www.reddit.com/r/ClaudeAI/comments/1s08r1c/karpathy_says_he_hasnt_written_a_line_of_code/"&gt;Reddit thread&lt;/a&gt; sparked a discussion about what some are calling "AI psychosis" or "cyber psychosis": Andrej Karpathy reportedly went from 80% writing his own code to 0%, spending 16 hours a day directing AI agents. Garry Tan described similar feelings of burning through 4 hours of sleep, unable to stop building.&lt;/p&gt;
&lt;p&gt;The debate that followed went beyond executives; it revealed a community-wide phenomenon: people running multiple Claude Code sessions in parallel, hitting rate limits daily, feeling like idle tokens were wasted tokens.&lt;/p&gt;
&lt;p&gt;The consensus across replies was clear: &lt;em&gt;AI psychosis&lt;/em&gt; is real, but it's less about excitement and more about a draining, addictive pressure to constantly build. The fear is missing out on the next big thing; it's the ground shifting beneath our feet, and stopping means getting left behind.&lt;/p&gt;
&lt;p&gt;Here's what I've found hard to articulate: AI tools have expanded possibility faster than ever; their real danger lies in how they collapse our attention span. Before AI, we operated on a fairly flat productivity curve: more effort meant more output, slowly but sustainably. Now we're running on a different kind of curve altogether—one that's getting steeper in both directions.&lt;/p&gt;
&lt;h2 id="the-accelerating-landscape-of-possibility"&gt;The Accelerating Landscape of Possibility&lt;/h2&gt;&lt;p&gt;In his book &lt;a href="https://libro.fm/audiobooks/9781101403860-where-good-ideas-come-from"&gt;Where Good Ideas Come From&lt;/a&gt;, Steven Johnson described what he called the &lt;em&gt;adjacent possible&lt;/em&gt;: the set of next-step ideas that are just beyond our current reality but still reachable. At any moment, only a limited set of next moves are accessible.&lt;/p&gt;
&lt;p&gt;Here's what makes this concept critical for understanding our current moment: as you explore the adjacent possible (through moves that seem natural in the moment) the boundary itself expands. Each discovery opens doors that weren't accessible before, creating an accelerating landscape where what's possible keeps growing faster and faster.&lt;/p&gt;
&lt;p&gt;I've seen this pattern play out: It explains why breakthroughs often happen when they do: the preconditions have finally assembled, and only then can you see the next move. But in an accelerating landscape, those preconditions assemble more quickly, and with them, the next set of possibilities.&lt;/p&gt;
&lt;h2 id="why-ai-feels-different-now"&gt;Why AI Feels Different Now&lt;/h2&gt;&lt;p&gt;Before AI, the adjacent possible was bounded by what a single person could manually assemble: write code, test it, debug it, repeat. The feedback loop, prompt, think, interpret, iterate, took time.&lt;/p&gt;
&lt;p&gt;AI tools changed that calculus. They transformed the feedback loop, making it exponentially faster. Each new capability doesn't just add possibility; it reconfigures what's adjacent.&lt;/p&gt;
&lt;p&gt;Once you see Claude Code as an &lt;em&gt;idea multiplier&lt;/em&gt;, the pattern is clear. The Garry Tan/Karpathy effect kicks in: possibility grows faster than effort.&lt;/p&gt;
&lt;h2 id="here-s-why-it-feels-different"&gt;Here's Why It Feels Different&lt;/h2&gt;&lt;p&gt;This is where it gets subtle: AI has shifted the inverted U curve of productivity and changed its shape.&lt;/p&gt;
&lt;p&gt;Here's what that looks like, the gray curve shows pre-AI productivity, and the red curve shows post-AI:&lt;/p&gt;
&lt;svg width="500" height="300" viewBox="0 0 500 300" xmlns="http://www.w3.org/2000/svg"&gt;
  &lt;!-- Axes --&gt;
  &lt;line x1="50" y1="250" x2="450" y2="250" stroke="#333" stroke-width="2"/&gt;
  &lt;line x1="50" y1="50" x2="50" y2="250" stroke="#333" stroke-width="2"/&gt;

  &lt;!-- Labels --&gt;
  &lt;text x="450" y="270" font-size="14" fill="#333"&gt;Effort&lt;/text&gt;
  &lt;text x="25" y="45" font-size="14" fill="#333"&gt;Output&lt;/text&gt;

  &lt;!-- Pre-AI curve (flatter, broader) --&gt;
  &lt;path d="M 80 230 Q 150 220 250 200 T 420 230" stroke="#888" stroke-width="3" fill="none"/&gt;
  &lt;text x="410" y="210" font-size="12" fill="#666" text-anchor="start"&gt;Pre-AI&lt;/text&gt;

  &lt;!-- Post-AI curve (steeper, narrower) --&gt;
  &lt;path d="M 80 230 Q 150 40 320 230" stroke="#e74c3c" stroke-width="3" fill="none"/&gt;
  &lt;text x="110" y="150" font-size="12" fill="#e74c3c" text-anchor="end"&gt;Post-AI&lt;/text&gt;
&lt;/svg&gt;&lt;p&gt;&lt;strong&gt;Pre-AI: flatter curve&lt;/strong&gt;. Effort mapped to output roughly proportionally. You could invest effort and see gains compound slowly, sustainably, with exhaustion coming only after sustained periods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post-AI: taller, narrower curve&lt;/strong&gt;. Less effort gets you more output initially; the left slope is steeper, giving an astonishing return on initial investment. The right-side drop-off is sharper—exhaustion hits earlier, harder.&lt;/p&gt;
&lt;p&gt;The peak represents the optimal effort level where output maximizes. Beyond that point, additional effort produces diminishing returns and exhaustion sets in faster than you can recover.&lt;/p&gt;
&lt;p&gt;I wrote about this in a previous post on closing air gaps: the problem is that things are faster; it's that the &lt;em&gt;gaps&lt;/em&gt; between tasks: those moments where attention bleeds out, are more expensive when everything moves faster.&lt;/p&gt;
&lt;h2 id="our-brain-is-the-bottleneck-now"&gt;Our Brain Is the Bottleneck Now&lt;/h2&gt;&lt;p&gt;AI gives us 10X faster feedback loops: code spits out, prompts happen in seconds. Neural processing remains capped at biological speeds.&lt;/p&gt;
&lt;p&gt;When loop cadence exceeds brain throughput, the cognitive queue overflows. Working memory saturates. Attention bleeds out. Exhaustion sets in, from the constant context switching, the constant need to &lt;em&gt;scan&lt;/em&gt; multiple sessions for "what was done."&lt;/p&gt;
&lt;p&gt;The optimal zone is when loop speed equals brain processing speed, not running tools as fast as possible.&lt;/p&gt;
&lt;h3 id="two-calibration-strategies"&gt;Two Calibration Strategies&lt;/h3&gt;&lt;p&gt;Here's what I believe we need to be able to do in order to calibrate properly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First, tighten feedback loops.&lt;/strong&gt; The goal is closing the loop properly on one task before opening another—I used to think juggling multiple Claude Code sessions was productivity; turns out it's just context switching masquerading as output. The trick is simple: run one session, close the loop completely, review what you got, then decide if your next move should be a new loop or something else entirely. Fast loops create natural pacing, which means you don't need to check tabs constantly because each cycle finishes with closure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second, build queue and notification systems—especially when you genuinely need multitasking.&lt;/strong&gt; Most of us reach for multiple agents because we're solving the wrong problem: juggling open sessions creates overhead your tools cannot absorb. The Kanban approach works well: externalize context switching onto a board where agents update their status automatically. The system notifies you only when human judgment is required, so your job shifts from scanning to assessing—much lower overhead for neural throughput.&lt;/p&gt;
&lt;h2 id="calibration-is-not-optimization"&gt;Calibration Is Not Optimization&lt;/h2&gt;&lt;p&gt;This is the crucial distinction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pre-AI&lt;/strong&gt;, you could coast for years on the left slope. Effort and reward grew linearly, so you just needed to work consistently.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post-AI&lt;/strong&gt;, the left slope is steeper, so you ascend faster; but also fall faster. AI-assisted tools don't eliminate the inverted U curve; they sharpen its peak and drop-off.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$\frac{d(\text{output})}{d(\text{effort})}$ is higher initially (good)&lt;/li&gt;
&lt;li&gt;$\frac{d^2(\text{output})}{d(\text{effort}^2)}$ is more negative; the curve drops off faster (bad)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Calibration recognizes that more effort does not equal more output. Instead, there exists an optimal effort level where output peaks. Beyond that point, additional effort produces diminishing returns, and rest becomes the superior strategy.&lt;/p&gt;
&lt;h2 id="what-calibration-actually-looks-like"&gt;What Calibration Actually Looks Like&lt;/h2&gt;&lt;p&gt;The bottleneck is neural throughput—not tokens or API calls.&lt;/p&gt;
&lt;p&gt;Ask yourself: Is my feedback loop faster than I can process?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If yes: slow down, close loops properly, resist the urge to open more tabs&lt;/li&gt;
&lt;li&gt;If no: optimize the tool, not your attention (this is where most of us are wrong)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Practical heuristic&lt;/strong&gt;: When you start scanning multiple AI threads for "what was done," you've exceeded your bandwidth. That's the signal that your loop cadence outruns your neural capacity.&lt;/p&gt;
&lt;p&gt;Ask, "What is the maximum rate at which my brain can consume and act on output?" This question defines calibration—when work aligns with your neural capacity.&lt;/p&gt;
&lt;h2 id="calibration-something-you-do-daily-not-something-you-learn-once"&gt;Calibration: Something You Do Daily, Not Something You Learn Once&lt;/h2&gt;&lt;p&gt;AI has revealed our limits: the inverted U curve has become more visible, accelerated. Our brain is the rate limiter, and rightly so! We now need to learn to brake before you hit the wall.&lt;/p&gt;
&lt;p&gt;The tools are powerful, but they don't change human neurology. No amount of prompt engineering can compress the time it takes for our brains to reason about things. If we try to go beyond our natural limits, dangerous things happen.&lt;/p&gt;
&lt;p&gt;Calibration is the new baseline practice: a discipline you maintain daily, adjusting your loop cadence to match neural throughput, closing gaps before they become exhaustions.&lt;/p&gt;
&lt;p&gt;The goal is sustainable access to the adjacent possible, rather than simply 10X-ing our output. And the goal is to do it for the long run, not just this week.&lt;/p&gt;
&lt;h3 id="what-this-looks-like-in-practice"&gt;What This Looks Like in Practice&lt;/h3&gt;&lt;p&gt;The goal is to work at the optimal effort level where output peaks. Beyond that point, additional effort produces diminishing returns and exhaustion sets in.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The default workflow:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One task at a time&lt;/strong&gt;: Start with &lt;em&gt;one&lt;/em&gt; focused task. Close the loop completely before moving to the next.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Close the loop fully&lt;/strong&gt;: Review output, make notes, then decide on the next step only when you're ready.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Eliminate scanning&lt;/strong&gt;: If you catch yourself flipping between tabs or sessions to check status, that's your signal that you've exceeded your neural throughput.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;When multiple agents must run simultaneously:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Use a Kanban board&lt;/strong&gt;: A visual task queue where agents update their status automatically.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agents update the board, not you&lt;/strong&gt;: The system should notify only when human intervention is needed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You check the board, not the sessions&lt;/strong&gt;: Remove the need to scan through terminal windows or tool interfaces.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You synchronize to stay on the left slope of your productivity curve, where effort yields return without triggering burnout. When tools run faster than your brain can process status updates, exhaustion sets in.&lt;/p&gt;
&lt;p&gt;The rhythm shifts from "hustle harder" to &lt;em&gt;synchronize&lt;/em&gt;—matching your brain's processing speed to the tools' output rate. When the signal outruns neural processing, interference replaces insight. You tune the system to keep your brain in phase on the left slope of your productivity curve, where effort produces sustainable output.&lt;/p&gt;
</content></entry><entry><title>Undoing AI vibe-coded slop with AI</title><link href="https://ericmjl.github.io/blog/2026/3/29/undoing-ai-vibe-coded-slop-with-ai/" rel="alternate"/><updated>2026-03-29T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:e7094c8e-a057-3915-aeaf-6efb27208eec</id><content type="html">&lt;p&gt;I want to tell you about canvas-chat, a project I built with heavy AI assistance. It's a visual, non-linear chat interface where conversations are nodes on an infinite canvas — think branching, merging, and exploring topics as a directed acyclic graph.&lt;/p&gt;
&lt;p&gt;The first commit landed on December 28, 2025. By December 30, it had sessions, matrix evaluation tables, web search, node tagging, and BM25 keyword search. The AI moved &lt;em&gt;fast&lt;/em&gt;. Bugs got fixed in the next commit. Features piled in like tetris blocks stacking up.&lt;/p&gt;
&lt;p&gt;And yes, it was a mess.&lt;/p&gt;
&lt;p&gt;Here's the thing though: the mess was &lt;em&gt;recoverable&lt;/em&gt;. Not because the AI got better (it didn't; not really), but because I had battle-tested convictions on how the thing &lt;em&gt;ought&lt;/em&gt; to be architected. And those convictions came from years of shipping software, watching architectures crumble, and learning what holds up.&lt;/p&gt;
&lt;p&gt;This is the story of how we went from a jumbled 8,500-line &lt;code&gt;app.js&lt;/code&gt; to a clean plugin architecture — and why you need battle-tested convictions to make that happen.&lt;/p&gt;
&lt;h2 id="the-initial-state"&gt;The Initial State&lt;/h2&gt;&lt;p&gt;The first commit wasn't actually &lt;em&gt;bad&lt;/em&gt;. The project had clean separation from day one:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;canvas.js&lt;/code&gt; — SVG pan/zoom/rendering&lt;/li&gt;
&lt;li&gt;&lt;code&gt;graph.js&lt;/code&gt; — DAG data structure&lt;/li&gt;
&lt;li&gt;&lt;code&gt;chat.js&lt;/code&gt; — LLM API + SSE streaming&lt;/li&gt;
&lt;li&gt;&lt;code&gt;storage.js&lt;/code&gt; — IndexedDB persistence&lt;/li&gt;
&lt;li&gt;&lt;code&gt;app.py&lt;/code&gt; — FastAPI backend&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But &lt;code&gt;app.js&lt;/code&gt; was already ~8,500 lines of everything else. Every slash command handler, every modal, every feature logic — all tangled together. Want to add a new feature? You'd grep around in that monolith, hope you found the right spot, and pray you didn't break anything.&lt;/p&gt;
&lt;p&gt;The AI could add features to this mess. It could add a &lt;code&gt;/matrix&lt;/code&gt; command in a few prompts. It could add &lt;code&gt;/search&lt;/code&gt; with Exa integration. But it couldn't see the &lt;em&gt;structure&lt;/em&gt; — the latent architecture that would make the whole thing maintainable.&lt;/p&gt;
&lt;h2 id="the-first-wave"&gt;The First Wave&lt;/h2&gt;&lt;p&gt;The refactoring started with the purest code — functions with no dependencies:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;What Got Extracted&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Jan 4&lt;/td&gt;
&lt;td&gt;&lt;code&gt;layout.js&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Overlap detection is pure math&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jan 5&lt;/td&gt;
&lt;td&gt;&lt;code&gt;highlight-utils.js&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Text selection is isolated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jan 7&lt;/td&gt;
&lt;td&gt;Feature modules&lt;/td&gt;
&lt;td&gt;&lt;code&gt;flashcards.js&lt;/code&gt;, &lt;code&gt;committee.js&lt;/code&gt;, &lt;code&gt;matrix.js&lt;/code&gt;, &lt;code&gt;factcheck.js&lt;/code&gt;, &lt;code&gt;research.js&lt;/code&gt;, &lt;code&gt;code.js&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jan 10&lt;/td&gt;
&lt;td&gt;Core infrastructure&lt;/td&gt;
&lt;td&gt;&lt;code&gt;undo-manager.js&lt;/code&gt;, &lt;code&gt;modal-manager.js&lt;/code&gt;, &lt;code&gt;slash-command-menu.js&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This reduced &lt;code&gt;app.js&lt;/code&gt; from ~8,500 to ~5,500 lines. But these were still just &lt;em&gt;file splits&lt;/em&gt;. The code worked better, but there was no &lt;em&gt;system&lt;/em&gt; binding it together.&lt;/p&gt;
&lt;p&gt;The AI did this part reasonably well — when I asked "extract this function to a separate module," it could do it. But it never suggested "we should extract this" on its own. It needed direction.&lt;/p&gt;
&lt;h2 id="the-pivotal-moment"&gt;The Pivotal Moment&lt;/h2&gt;&lt;p&gt;This was the architectural leap. I asked the AI to create a plugin system, and it delivered — but only because I knew what a plugin system &lt;em&gt;should&lt;/em&gt; look like.&lt;/p&gt;
&lt;p&gt;We ended up with a &lt;strong&gt;three-level plugin architecture&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Level 1: Custom Node Types&lt;/strong&gt; — Node protocols define rendering via a &lt;code&gt;BaseNode&lt;/code&gt; class. Each node type can override &lt;code&gt;renderContent()&lt;/code&gt;, &lt;code&gt;getActions()&lt;/code&gt;, &lt;code&gt;getSummaryText()&lt;/code&gt;, and more. Registered in &lt;code&gt;node-registry.js&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Level 2: Feature Plugins&lt;/strong&gt; — Extend a &lt;code&gt;FeaturePlugin&lt;/code&gt; base class. Get &lt;code&gt;AppContext&lt;/code&gt; via dependency injection (graph, canvas, chat, storage, modalManager, streamingManager). Define slash commands via &lt;code&gt;getSlashCommands()&lt;/code&gt;. Lifecycle hooks: &lt;code&gt;onLoad()&lt;/code&gt;, &lt;code&gt;onUnload()&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Level 3: Extension Hooks&lt;/strong&gt; — Subscribe to events. &lt;code&gt;CancellableEvent&lt;/code&gt; can block actions. Event names like &lt;code&gt;command:before&lt;/code&gt;, &lt;code&gt;node:created&lt;/code&gt;, &lt;code&gt;node:deleted&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The key files created:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;feature-plugin.js&lt;/code&gt; — FeaturePlugin + AppContext&lt;/li&gt;
&lt;li&gt;&lt;code&gt;feature-registry.js&lt;/code&gt; — Slash command routing with priority (BUILTIN &amp;gt; OFFICIAL &amp;gt; COMMUNITY)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;plugin-events.js&lt;/code&gt; — CanvasEvent, CancellableEvent&lt;/li&gt;
&lt;li&gt;&lt;code&gt;node-registry.js&lt;/code&gt; — Node type registration&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is where the architecture became a real system. And it only happened because I knew what I wanted.&lt;/p&gt;
&lt;h2 id="backend-pluginification-late-january-2026"&gt;Backend Pluginification (Late January 2026)&lt;/h2&gt;&lt;p&gt;The same pattern reached the Python side:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;pptx_endpoints.py&lt;/code&gt; — PowerPoint handling&lt;/li&gt;
&lt;li&gt;&lt;code&gt;ddg_endpoints.py&lt;/code&gt; — DuckDuckGo search&lt;/li&gt;
&lt;li&gt;&lt;code&gt;code_handler.py&lt;/code&gt; — Python code execution&lt;/li&gt;
&lt;li&gt;&lt;code&gt;matrix_handler.py&lt;/code&gt; — Matrix cell filling&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each follows a &lt;code&gt;register_endpoints(app)&lt;/code&gt; pattern, loaded dynamically via &lt;code&gt;importlib&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="the-testing-safety-net"&gt;The Testing Safety Net&lt;/h2&gt;&lt;p&gt;By late January, the plugin architecture was in place. Features were decoupled. The code was cleaner. And then GLM-4.5 started dropping curly braces.&lt;/p&gt;
&lt;p&gt;No, really. The AI would "fix" one thing and introduce a missing bracket somewhere else. Merge conflicts became minefields; features that worked yesterday stopped working today; not because of malice, but because the AI didn't understand the dependencies between modules. It was making elementary mistakes that a junior developer wouldn't make.&lt;/p&gt;
&lt;p&gt;On January 24, I added Cypress E2E tests. Out of spite, honestly. The first commit gave us &lt;code&gt;canvas_interactions.cy.js&lt;/code&gt;, &lt;code&gt;node_selection.cy.js&lt;/code&gt;, and &lt;code&gt;note_node.cy.js&lt;/code&gt; - three tests that told us whether the canvas still worked.&lt;/p&gt;
&lt;p&gt;These tests caught the regressions the AI kept introducing. More importantly, they let me verify changes faster. Instead of manually testing every feature after each AI session, I could run the test suite and know whether things still worked.&lt;/p&gt;
&lt;p&gt;The plugin architecture made the code testable. The tests caught what the AI broke.&lt;/p&gt;
&lt;h2 id="the-numbers"&gt;The Numbers&lt;/h2&gt;&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;app.js size&lt;/th&gt;
&lt;th&gt;Modules&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Initial (Dec 2025)&lt;/td&gt;
&lt;td&gt;~8,500 lines&lt;/td&gt;
&lt;td&gt;5 files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After feature splits&lt;/td&gt;
&lt;td&gt;~5,500 lines&lt;/td&gt;
&lt;td&gt;11 files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After infrastructure&lt;/td&gt;
&lt;td&gt;~5,400 lines&lt;/td&gt;
&lt;td&gt;15 files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;After plugin migration&lt;/td&gt;
&lt;td&gt;~5,400 lines&lt;/td&gt;
&lt;td&gt;25+ files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Today&lt;/td&gt;
&lt;td&gt;~4,700 lines&lt;/td&gt;
&lt;td&gt;35+ modules&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="the-bigger-lesson"&gt;The Bigger Lesson&lt;/h2&gt;&lt;p&gt;Here's what I learned from this process:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The AI can execute architecture, but it can't design it.&lt;/strong&gt; It can split files when asked. It can implement a plugin system from a spec. But it won't look at a 8,500-line &lt;code&gt;app.js&lt;/code&gt; and say "this should be a plugin system."&lt;/p&gt;
&lt;p&gt;That vision; that &lt;em&gt;opinion&lt;/em&gt;; comes from somewhere else. It comes from:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Seeing architectures fail&lt;/strong&gt; - Knowing the pain of tangled code, merged conflicts, feature creep&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Seeing architectures succeed&lt;/strong&gt; - Knowing what maintainable code feels like after years of shipping&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reading, studying, internalizing&lt;/strong&gt; - Design patterns, architectural styles, tradeoffs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Making mistakes&lt;/strong&gt; - Building the wrong abstraction once so you recognize it next time&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I didn't arrive at "we need a three-level plugin architecture" out of nowhere. It came from discussing tradeoffs with the AI; asking "what if we did it this way?" and "what are the tradeoffs of that approach?"; and applying my best judgment to the options. The AI could explain the pros and cons of different approaches, but I had to pick which tradeoffs I was willing to accept.&lt;/p&gt;
&lt;p&gt;The AI didn't teach me this. &lt;em&gt;Experience&lt;/em&gt; taught me this.&lt;/p&gt;
&lt;h2 id="what-this-means-for-the-future"&gt;What This Means for the Future&lt;/h2&gt;&lt;p&gt;Here's where it gets interesting. Because we built this modular foundation, I can now swap out the rendering layer. The canvas is currently raw SVG — and I want to move to Svelte Flow. The plugin system I built makes this possible:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Features don't depend on &lt;code&gt;app.js&lt;/code&gt; internals; they use &lt;code&gt;AppContext&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Canvas is isolated in &lt;code&gt;canvas.js&lt;/code&gt;; swapping to Svelte Flow means replacing that layer&lt;/li&gt;
&lt;li&gt;Node protocols define behavior; Svelte Flow nodes can use the same protocol pattern&lt;/li&gt;
&lt;li&gt;Event system is framework-agnostic&lt;/li&gt;
&lt;li&gt;Dependency injection provides graph, canvas, chat, storage; these can be re-provided to Svelte Flow components&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The abstraction layer we built; FeaturePlugin + AppContext + EventSystem; separates &lt;em&gt;what&lt;/em&gt; features do from &lt;em&gt;how&lt;/em&gt; they're rendered. That's what makes Svelte Flow viable as a drop-in replacement.&lt;/p&gt;
&lt;h2 id="closing-thoughts"&gt;Closing Thoughts&lt;/h2&gt;&lt;p&gt;You can undo AI vibe-coded slop. It's possible. But it requires &lt;em&gt;you&lt;/em&gt; to have battle-tested convictions on how the thing ought to be.&lt;/p&gt;
&lt;p&gt;The AI is an incredible executor. It can refactor, extract, implement. But the vision? That stays human. And that vision comes from battle-tested experience; from having seen enough codebases to know what works and what collapses under its own weight.&lt;/p&gt;
&lt;p&gt;So if you're working with AI coding assistants: don't expect them to architect for you. Tell them what to build. Give them the structure. Then let them do the implementation.&lt;/p&gt;
&lt;p&gt;That's how you get from a jumbled mess to something you can actually maintain.&lt;/p&gt;
</content></entry><entry><title>Creative mentorship strategies for career growth in challenging times</title><link href="https://ericmjl.github.io/blog/2026/3/25/creative-mentorship-strategies-for-career-growth-in-challenging-times/" rel="alternate"/><updated>2026-03-25T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:9e5abdb7-d7d1-32b6-9593-332abe1cdbb5</id><content type="html">&lt;h2 id="the-problem-with-lean-times"&gt;The problem with lean times&lt;/h2&gt;&lt;p&gt;When the economy tightens, formal development opportunities are usually the first things to go. Co-ops get paused, training budgets shrink, and headcount freezes make it harder to bring in fresh talent. But the need to develop mentorship, coaching, and leadership skills doesn't disappear just because the budget did.&lt;/p&gt;
&lt;p&gt;So the question becomes: how do you get creative? How do you find opportunities to grow as a mentor and leader without requiring the company to spend additional money?&lt;/p&gt;
&lt;h2 id="you-already-have-something-to-offer"&gt;You already have something to offer&lt;/h2&gt;&lt;p&gt;The answer is closer than you think. Even when budgets are frozen, you still have three things worth sharing: your judgment, your skills, and your network.&lt;/p&gt;
&lt;p&gt;Your judgment is what experience actually gives you: not just knowing things, but knowing which approach to take, which tradeoffs matter, and when to push versus when to hold back. Your skills are the technical foundation that lets you coach and mentor beyond your own team, helping others onboard to the tools and practices you work with. And your network is the set of connections you can activate to create opportunities for others, whether that means knowing the right organizer, connecting a speaker to an audience, or simply inviting people into the same room.&lt;/p&gt;
&lt;p&gt;The people who want to learn from you are already in your neighborhood. If you are doing an excellent job, you will find individuals who are eager to understand how you achieve your results. That is where your opportunity lies.&lt;/p&gt;
&lt;h2 id="five-strategies-that-have-worked-for-me"&gt;Five strategies that have worked for me&lt;/h2&gt;&lt;p&gt;I want to share some concrete strategies that have worked at the two companies I have been with, Novartis and Moderna. Some of these I have actively advocated for. My intent is not to boast but to provide pragmatic suggestions based on my own experiences.&lt;/p&gt;
&lt;h3 id="coach-others-one-on-one"&gt;Coach others one-on-one&lt;/h3&gt;&lt;p&gt;Coaching others is a great way to build your reputation within the organization. When you teach someone how to accomplish a task effectively, you become valuable to them. More importantly, you demonstrate your value to a broader audience. Within an organization, you want a group of people who find your skills genuinely worthwhile.&lt;/p&gt;
&lt;h3 id="present-at-internal-guilds-and-birds-of-a-feather-events"&gt;Present at internal guilds and "birds of a feather" events&lt;/h3&gt;&lt;p&gt;At Moderna's Digital organization, we have "Guilds", the Data Science Guild being one, with three meetings per month for the guild. When I was at Novartis' Research org, we had the Computational Research Community. Both served as outlets for talks and annual gatherings. The key is to be in a position where you can give a talk about something valuable to others, and they would willingly spend an hour listening to you. If that happens, you have created another mentorship opportunity for yourself.&lt;/p&gt;
&lt;h3 id="organize-communities-of-practice"&gt;Organize communities of practice&lt;/h3&gt;&lt;p&gt;I have seen this happen when someone builds a tool that others use and then creates a community around that tool. It can be as simple as a Microsoft Teams group chat. You don't need anything more sophisticated than that. Just gather people who use the tool and facilitate discussions. Being a leader in that group chat is a real way to hone your leadership skills across the organization. My colleague &lt;a href="https://www.linkedin.com/in/albert-lam/"&gt;Albert Lam&lt;/a&gt; built a significant portion of the Python packages that are used by LLM builders, and put together communities of practice around that precisely in the form of MS Teams group chats.&lt;/p&gt;
&lt;p&gt;Another example is the community of practice around documentation - primarily expressed as the internal &lt;a href="../../../../2024/6/30/two-years-of-docathons-insights-and-lessons-learned/"&gt;docathons&lt;/a&gt; that we run. My teammate &lt;a href="https://jackievaleri.github.io/"&gt;Jackie Valeri&lt;/a&gt;, as well as two other colleagues &lt;a href="https://www.linkedin.com/in/simreen-kaur/"&gt;Simreen Kaur&lt;/a&gt; and &lt;a href="https://www.linkedin.com/in/saakshisdonthi/"&gt;Saakshi Shamanth Donthi&lt;/a&gt; help coordinate and organize the logistics while also being point contacts for other folks participating in the docathon.&lt;/p&gt;
&lt;h3 id="host-informal-coffee-hours"&gt;Host informal coffee hours&lt;/h3&gt;&lt;p&gt;My teammate, &lt;a href="https://mfaits.github.io/"&gt;Michelle Faits&lt;/a&gt;, took the initiative to host coffee hours within the company. These serve as informal outlets for people to present their work, and they are great because they are relaxed and authentic. As her manager, I try to find speakers to contribute and support her efforts. Kudos to her for initiating this.&lt;/p&gt;
&lt;h3 id="host-or-support-external-meetups"&gt;Host or support external meetups&lt;/h3&gt;&lt;p&gt;We also host the PyData Boston/Cambridge monthly meetup at Moderna. Not every month is held at our location, but since I know the organizer &lt;a href="https://benbatorsky.com/"&gt;Ben Batorsky&lt;/a&gt;, back in 2025, I offered my time to book a room; we simply provide the space. More recently, Jackie has taken the lead in this. By taking charge and inviting others to network, we create opportunities for people to grow in their careers without any budget requirement.&lt;/p&gt;
&lt;h2 id="advice-for-managers"&gt;Advice for managers&lt;/h2&gt;&lt;p&gt;If you are a manager, recognize that there will be projects and efforts that need leadership beyond individual contributions. Helping others improve their skills and providing them with opportunities to lead is part of the job, especially when formal avenues are limited.&lt;/p&gt;
&lt;p&gt;Make sure you are aware of these kinds of initiatives among your reports, and ensure they don't conflict with core responsibilities. When someone can demonstrate that they manage these additional activities while maintaining their primary work, that forms a strong case for their expanded capabilities.&lt;/p&gt;
&lt;p&gt;We should expand our imagination beyond just climbing the career ladder, attaining higher status, or managing other people, which sometimes unfortunately spills over into controlling others at work. Growth comes in many forms, and we can find meaningful ways to develop without waiting for formal promotions or titles.&lt;/p&gt;
&lt;h2 id="the-core-principles"&gt;The core principles&lt;/h2&gt;&lt;p&gt;Mentoring is about sharing your judgment, providing opportunities for others to share theirs, facilitating networking, and helping others grow. If we limit our understanding of growth and development to a narrow definition (only formal programs, only budgeted activities, only additional formal assignments), we constrain our imagination as managers.&lt;/p&gt;
&lt;p&gt;Don't be confined to a singular vision of what it means to be a good leader or manager. Embrace the diverse talents and varying stages of abilities within your team. Encourage your team members to step outside their comfort zones and provide them with opportunities to grow.&lt;/p&gt;
&lt;p&gt;While good, you do not necessarily need an internal university to foster a learning culture; your environment is where you learn and grow. We have more autonomy and agency than we might realize!&lt;/p&gt;
</content></entry><entry><title>Closing air gaps</title><link href="https://ericmjl.github.io/blog/2026/3/15/closing-air-gaps/" rel="alternate"/><updated>2026-03-15T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:739d4d76-ce27-3a06-a50b-cffe98d01bbb</id><content type="html">&lt;p&gt;I owe this term to my colleague &lt;a href="https://www.linkedin.com/in/wenhao-liu-1b85177/"&gt;Wenhao Liu&lt;/a&gt;. He was the first one I saw at work who clearly articulated about air gaps and how they relate to building agents for work.&lt;/p&gt;
&lt;p&gt;So what exactly is an air gap? It is any point in a business or scientific process where a human has to intervene and perform manual work before a digital system can continue. The system cannot go end to end on its own; the human is the bridge.&lt;/p&gt;
&lt;p&gt;Air gaps are everywhere. Here are a few examples: A laboratory machine exports a file to a local hard disk. A human copies that file and pastes it into an S3 bucket. That handoff is an air gap. The system stops at the hard disk and waits for a person.&lt;/p&gt;
&lt;p&gt;In wet lab science, air gaps take physical form. A scientist designs an experiment, walks into the lab, and executes it by hand. In this state, the company/team/org has an air gap in its scientific process. No robotic system can take over from design to execution.&lt;/p&gt;
&lt;p&gt;Both of these examples share a common pattern. The definition I have settled on is this: an air gap is any place where rote manual work is performed by a human that could have been done by a computer. (Robots are computers with sensors and actuators for the physical world.)&lt;/p&gt;
&lt;h2 id="why-air-gaps-matter"&gt;Why air gaps matter&lt;/h2&gt;&lt;p&gt;Think of your processes as pipes. Air gaps are bubbles trapped in those pipes. They slow the flow. They disrupt continuity. They introduce delays and errors.&lt;/p&gt;
&lt;p&gt;The costs compound over time. A five-minute manual handoff, repeated daily across a team of twenty, adds up to real hours. A week-long delay because someone was on vacation and could not move the file. A transcription error because a human typed a number wrong. A lost opportunity because the data sat in a local folder instead of flowing into the analysis pipeline.&lt;/p&gt;
&lt;p&gt;Air gaps also create cognitive overhead. Every time a human has to remember to perform a manual step, that is mental bandwidth not spent on creative work. The air gap is a tax on attention.&lt;/p&gt;
&lt;p&gt;Now, a clarification. The goal here is not to eliminate humans from everything. The goal is to eliminate humans from the rote and routine. Creative work, judgment, and decision making stay with us. Copying files, transferring plates, and typing data into forms do not.&lt;/p&gt;
&lt;h2 id="how-to-find-air-gaps"&gt;How to find air gaps&lt;/h2&gt;&lt;p&gt;You cannot close an air gap until you see it. And you cannot see it until you sit down and map out exactly how your process works. In other words, &lt;strong&gt;process mapping&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This mapping exercise is the unsexy work that precedes automation. Most teams skip it. They jump to solutions before understanding the problem. But you need the map.&lt;/p&gt;
&lt;p&gt;Here is how to do it.&lt;/p&gt;
&lt;p&gt;Pick one process. It could be a data pipeline, a lab workflow, or a business approval chain. Walk through it step by step. Ask these questions at each step:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Where is a human manually copying and pasting?&lt;/li&gt;
&lt;li&gt;Where is a human manually entering data?&lt;/li&gt;
&lt;li&gt;Where is a human dragging and dropping files?&lt;/li&gt;
&lt;li&gt;Where is a human making a decision that follows a fixed rule?&lt;/li&gt;
&lt;li&gt;Where is a human waiting for another human to take action?&lt;/li&gt;
&lt;li&gt;Where is a human physically moving something from one place to another?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each answer points to a potential air gap.&lt;/p&gt;
&lt;p&gt;Write it down. Draw it out. Make the process visible. Once the map exists, the air gaps reveal themselves. The next step is prioritization. Which air gaps cause the most pain? Which ones are easiest to close? Start there.&lt;/p&gt;
&lt;h2 id="air-gaps-in-the-wild"&gt;Air gaps in the wild&lt;/h2&gt;&lt;p&gt;Enough abstraction. Let me show you what air gaps look like in practice, from my own work and from the broader landscape.&lt;/p&gt;
&lt;h3 id="file-schlepping-in-the-lab"&gt;File schlepping in the lab&lt;/h3&gt;&lt;p&gt;A sequencing machine finishes a run. It writes the data to a local drive. A technician notices the run is complete, navigates to the folder, selects the files, copies them, navigates to the shared storage system, and pastes. Minutes pass. Sometimes hours, if the technician is busy.&lt;/p&gt;
&lt;p&gt;This is an air gap. The machine knows when the run finishes. The machine has network access. The destination storage has an API. Automation can close this gap with a simple script that watches for new files and uploads them.&lt;/p&gt;
&lt;p&gt;The fix is not technically difficult; it is, conceptually, a &lt;code&gt;cron&lt;/code&gt; job with &lt;code&gt;rsync&lt;/code&gt;. What makes it hard is that the air gap is invisible until someone maps the process and asks why a human is doing this work.&lt;/p&gt;
&lt;h3 id="github-activity-tracking"&gt;GitHub activity tracking&lt;/h3&gt;&lt;p&gt;I used to manually check GitHub to track my daily work. I would open my profile page, scroll through recent commits, open the pull requests tab, check which ones I had opened or reviewed, and then type notes into a document. This took maybe ten minutes per day.&lt;/p&gt;
&lt;p&gt;Then I remembered the GitHub CLI exists. I also remembered that coding agents can run CLI commands.&lt;/p&gt;
&lt;p&gt;I built a skill that pulls four categories of activity automatically: my opened pull requests, pull requests I reviewed or commented on, my commits, and issues I created. The agent runs this skill as part of my daily sign-off routine. The air gap closed.&lt;/p&gt;
&lt;p&gt;The time savings are modest. But the mental overhead vanished. I no longer need to remember to check GitHub. The information flows to me.&lt;/p&gt;
&lt;h3 id="autonomous-laboratories"&gt;Autonomous laboratories&lt;/h3&gt;&lt;p&gt;The autonomous lab, sometimes called a lights-out lab, is the ultimate expression of closing air gaps. The vision is a laboratory that runs itself: experiments are designed, executed, analyzed, and iterated without human intervention.&lt;/p&gt;
&lt;p&gt;In practice, autonomous labs are full of micro air gaps. Each one must be identified and closed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plate transfers.&lt;/strong&gt; A protocol requires moving a plate from one instrument to another. Does a human do this? If so, that is an air gap. Robotic arms and conveyor systems can close it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Master reagent prep.&lt;/strong&gt; Someone mixes buffers and reagents by hand at the start of each week. Could a liquid handling robot do this instead? Probably. That is an air gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;File movement.&lt;/strong&gt; Instruments write data locally. Humans move data to shared storage. This is the file schlepping problem again, repeated across every machine in the lab.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Standardized analyses.&lt;/strong&gt; As a data scientist, this one is close to my heart. Most labs have a set of standard analyses they run on every dataset. Quality control plots, basic statistics, alignment checks. A human opens a notebook, loads the data, runs the cells, and exports results.&lt;/p&gt;
&lt;p&gt;This is an air gap. Standardized analyses can be automated. They can also be made adaptable. An LLM-powered coding agent can take a standard analysis template and adjust parameters within a confined range of design choices. The human specifies the intent. The agent handles the implementation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Closing the loop.&lt;/strong&gt; The ultimate air gap in scientific research is the gap between analysis and experiment design. A human looks at results, draws conclusions, and designs the next experiment. What if the results could flow back into experiment design automatically? What if an agent could propose the next experiment based on what the data showed?&lt;/p&gt;
&lt;p&gt;This is the direction autonomous labs are moving. But getting there requires closing every air gap along the chain.&lt;/p&gt;
&lt;h2 id="the-blockers-imagination-and-skill"&gt;The blockers: imagination and skill&lt;/h2&gt;&lt;p&gt;I wrote about this in a &lt;a href="https://ericmjl.github.io/blog/2026/3/6/mastering-personal-knowledge-management-with-obsidian-and-ai/"&gt;previous post&lt;/a&gt;. The biggest blockers to closing air gaps are technical skill and imagination.&lt;/p&gt;
&lt;h3 id="imagination"&gt;Imagination&lt;/h3&gt;&lt;p&gt;If you cannot imagine a future state where your tedious work is performed by a coding agent, you will not see the possibility. The air gap remains invisible.&lt;/p&gt;
&lt;p&gt;This is a failure of imagination, not a failure of technology. The tools exist. The APIs exist. The agents exist. What is missing is the mental leap from "this is how we have always done it" to "this is how we could do it."&lt;/p&gt;
&lt;p&gt;Imagination grows from exposure. The more you see what is possible, the more you can imagine for your own work. Watch what other teams are doing. Read about automation in adjacent fields. Talk to people who have closed similar gaps.&lt;/p&gt;
&lt;h3 id="skill"&gt;Skill&lt;/h3&gt;&lt;p&gt;Imagination alone is not enough. You also need the skill to build the automation.&lt;/p&gt;
&lt;p&gt;That skill, at some level, means knowing how to program. It means understanding APIs, scripting, and how systems talk to each other. It means knowing that cron jobs and web hooks can be configured.&lt;/p&gt;
&lt;p&gt;The good news is that the barrier to entry is lower than ever. Coding agents can help you write the code. The skill you need is not deep software engineering. It is enough programming literacy to describe what you want and recognize whether the output is correct.&lt;/p&gt;
&lt;h3 id="programmatic-access"&gt;Programmatic access&lt;/h3&gt;&lt;p&gt;Sometimes the blocker is not you. Sometimes it is the system.&lt;/p&gt;
&lt;p&gt;Some organizations block programmatic access to their tools. It may be that a SaaS application has no API, or the API is disabled for security reasons, or that a legacy database has no query interface, only a web portal. A legacy laboratory information management system may require a human to click through screens instead of secured APIs.&lt;/p&gt;
&lt;p&gt;These are also air gaps. The system itself prevents automation.&lt;/p&gt;
&lt;p&gt;If the blocker is cybersecurity concerns, push for scoped, tracked access. You do not need full API permissions to close a specific air gap. You need enough access to move the data that belongs in your pipeline. That access can be limited, logged, and auditable.&lt;/p&gt;
&lt;p&gt;If the system truly has no programmatic interface, browser or desktop automation agents can close the gap. A headless browser can log in, navigate, click, and extract. It is slower and more fragile than an API, but it works.&lt;/p&gt;
&lt;h2 id="closing-air-gaps-with-agents"&gt;Closing air gaps with agents&lt;/h2&gt;&lt;p&gt;But here is the good news. The rise of coding agents changes the calculus for closing air gaps.&lt;/p&gt;
&lt;p&gt;Before, you needed a software engineer to write the automation script. Now, you can describe what you want in plain language and let the agent write the code.&lt;/p&gt;
&lt;p&gt;This does not mean you can ignore technical literacy. You still need to verify the output, debug when things break, and understand enough to specify the problem clearly. But the implementation barrier is lower.&lt;/p&gt;
&lt;p&gt;Browser agents extend this further. If a system has no API, a browser agent can act as the interface. Log in, click the buttons, extract the data, and feed it into your pipeline.&lt;/p&gt;
&lt;p&gt;The key insight is that agents are not a replacement for mapping your processes. They are a tool for closing the air gaps you have already identified. The mapping still matters. The imagination still matters. The skill to recognize whether the agent's output is correct still matters.&lt;/p&gt;
&lt;p&gt;What changes is the speed of iteration. You can try closing an air gap in an afternoon instead of a sprint. You can experiment with different approaches quickly. The feedback loop tightens.&lt;/p&gt;
&lt;p&gt;Here is the right question to ask when you find yourself being asked to do something repeatedly: "Can I remove this bottleneck for you? Is there a way I can make it so that you never have to ask me this again?" The goal is to build systems that use agents to not need human intervention as much as possible, not to create systems that depend on humans more and more. As one person put it, "the more I have asked myself that question, the more capable he has become." (&lt;a href="https://www.youtube.com/watch?v=nSBKCZQkmYw"&gt;source&lt;/a&gt;)&lt;/p&gt;
&lt;h2 id="the-principle"&gt;The principle&lt;/h2&gt;&lt;p&gt;The guiding principle is simple. To butcher the Biblical phrase, "Give unto robots what belongs to robots, and have humans do what humans can do." Or to paraphrase another person, use robots for the dull, dirty, and dangerous work.&lt;/p&gt;
&lt;p&gt;Keep the human in the loop for creative, judgment-heavy work. The rote and routine should flow through pipes without bubbles.&lt;/p&gt;
&lt;p&gt;Start mapping. Find the air gaps. Close them one by one. The compounding effect over time is enormous once micro-efficiencies become part of your work.&lt;/p&gt;
&lt;p&gt;What looks like a small efficiency gain today becomes a transformed process tomorrow. The lab that closed its file schlepping air gaps is one step closer to autonomous operation. The team that automated their daily reporting has mental bandwidth for harder problems.&lt;/p&gt;
&lt;p&gt;Close every air gap you can see!&lt;/p&gt;
</content></entry><entry><title>Agent skills are also human skills</title><link href="https://ericmjl.github.io/blog/2026/3/14/agent-skills-are-also-human-skills/" rel="alternate"/><updated>2026-03-14T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:3258471c-f953-3f42-92b1-712fa77e264b</id><content type="html">&lt;p&gt;Agent skills are great, but I've been thinking about this... skills alone aren't enough.&lt;/p&gt;
&lt;p&gt;I've been thinking about this while developing and using agent skills at home and at work. There's a distinction I've started to draw between two types of skills. Tool-specific skills document how to work with a particular tool or package. Those are fine, but really, pointing an agent at &lt;code&gt;llms.txt&lt;/code&gt; often works just as well. The more interesting category is workflow-specific skills, things that encode how you actually work, that string together multiple tools to &lt;strong&gt;get a job done&lt;/strong&gt; (Christensen).&lt;/p&gt;
&lt;p&gt;Workflow-specific skills are what I want to talk about here.&lt;/p&gt;
&lt;h2 id="a-concrete-example"&gt;A concrete example&lt;/h2&gt;&lt;p&gt;My daily sign-off skill, which I use at work, is a case in point. I use it to wrap up my day. When I sign off, I need two things: my meeting notes (which I paste into Obsidian throughout the day) and my GitHub activity (commits, PRs, comments, reviews). The skill handles the GitHub part by querying the GitHub CLI and formatting everything into my daily bullets template.&lt;/p&gt;
&lt;p&gt;But here's where it gets opinionated. My skill assumes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;You have the GitHub CLI installed&lt;/li&gt;
&lt;li&gt;You do PRs as part of your work (not all technical managers do)&lt;/li&gt;
&lt;li&gt;You write into a monthly file as your bullet journal, rather than having a single note per day.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That last point is opinionated. I don't have a single note per day. Instead, each month contains my collection of daily bullets. The motivation here is a line from the Zen of Python -- "flat is better than nested". On March 26, I have entries for that day inside the March file, rather than have a reference from the March file to March 26. This might not reflect your own preferences; you might prefer one note per day, or use a different structure entirely. But this is what my skill expects, and it's baked into how the skill works.&lt;/p&gt;
&lt;p&gt;If you want to use my daily sign-off skill, you're not just adopting the skill. You're adopting my way of working. You're inheriting my file structure, my tool preferences, my mental model for organizing information. The skill comes with implicit assumptions about how you work, what tools you use, and what your environment looks like.&lt;/p&gt;
&lt;h2 id="a-second-example-cutting-deeper"&gt;A second example, cutting deeper&lt;/h2&gt;&lt;p&gt;The daily sign-off is mostly about tool and structure preferences. But some skills go further — they encode a &lt;em&gt;philosophy&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;My scientific EDA skill is a good example. On the surface it looks like a set of technical rules: use &lt;code&gt;uv&lt;/code&gt; with PEP723 inline scripts, save plots as WebP (not PNG), organize each analysis session into a timestamped folder, keep an append-only &lt;code&gt;journal.md&lt;/code&gt;. But look at what those rules actually encode:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One step at a time, ask "why" before executing&lt;/strong&gt; — this isn't a technical constraint. It reflects a skepticism of agents that run ahead of the analyst. I believe good exploratory analysis is a dialogue, not a sprint.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Capture the research question before touching the data&lt;/strong&gt; — this reflects a conviction that context shapes what you should even be looking for. Data without a question is just noise.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Append-only journal&lt;/strong&gt; — this reflects a belief that good science is narrated, not just executed. The journal isn't a log file; it's a record of reasoning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;WebP over PNG&lt;/strong&gt; — a small but deliberate aesthetic and practical stance on file hygiene.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;uv + PEP723&lt;/strong&gt; — a specific bet on the Python toolchain that not everyone has made.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;None of these are neutral defaults. Each one is a choice that reflects how I think scientific work should be done. If you use my EDA skill but don't share that underlying philosophy, you'll find yourself fighting it. The one-step-at-a-time rule will feel like friction. The journaling requirement will feel like overhead. The skill isn't broken — it's just mine.&lt;/p&gt;
&lt;p&gt;This is a different kind of assumption from the daily sign-off. There, you're inheriting my tools and file layout. Here, you're inheriting my epistemology. That's harder to see, harder to document, and harder to transfer.&lt;/p&gt;
&lt;h2 id="what-this-means"&gt;What this means&lt;/h2&gt;&lt;p&gt;I call this &lt;em&gt;procedural context&lt;/em&gt;. A workflow-specific agent skill is more than documentation for the coding agent. It also implicitly encodes a person's systems and structures for working. Without documenting the procedural context, the skill can only be half-useful for another person.&lt;/p&gt;
&lt;p&gt;The two examples above hint at different layers of procedural context. There are at least three:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Tool dependencies&lt;/strong&gt; — what software needs to be installed (GitHub CLI, &lt;code&gt;uv&lt;/code&gt;, etc.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Organizational preferences&lt;/strong&gt; — how you structure files, folders, and notes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Epistemic preferences&lt;/strong&gt; — how you believe the work should actually proceed&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The third layer is the most invisible. It's also the most important, and the hardest to transfer. You can install a CLI tool in five minutes. Adopting someone else's philosophy of scientific analysis is a different ask entirely.&lt;/p&gt;
&lt;p&gt;Someone on Twitter put it well (I wish I could remember who, so I won't take credit): with agent skills, we finally found a way to get coders to write documentation. We'll document how we work if it means we can delegate that work to some{one/thing} else!&lt;/p&gt;
&lt;p&gt;At the end of the day, agent skills are just automation and documentation. We're automating away the minutiae, and I love that. But if your skill describes a workflow, you need to document the assumptions too. What are the dependencies? What tools need to be installed? What mental structures does the person need? What does the user need to know to verify the output is correct?&lt;/p&gt;
&lt;p&gt;Without that context, you can't evaluate whether the coding agent used the skill correctly -- and verification matters! You need to know what to look for when an LLM does work on your behalf.&lt;/p&gt;
&lt;h2 id="the-takeaway"&gt;The takeaway&lt;/h2&gt;&lt;p&gt;Agent skills implicitly involve human skills. If that's true, then agent skills are also for humans. They're not merely instructions for an agent. They're documentation of how someone accomplishes a job, with all the prerequisites and context needed to reproduce it.&lt;/p&gt;
&lt;p&gt;So when you write a workflow skill, think about the other people who might use it. Ask the skill-creator skill to include the dependencies, explain the environment, and describe what success looks like. The skill alone isn't enough. We have to teach the next person how to use it too.&lt;/p&gt;
</content></entry><entry><title>My weekend experiment making PyMC installable in a WASM environment</title><link href="https://ericmjl.github.io/blog/2026/3/8/my-weekend-experiment-pymc-wasm/" rel="alternate"/><updated>2026-03-08T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:bec96de2-ebd1-3052-a10c-136d1b83cc99</id><content type="html">&lt;p&gt;This past weekend, I found myself revisiting a blog post from PyMC Labs titled &lt;a href="https://www.pymc-labs.com/blog-posts/pymc-in-browser"&gt;"Running PyMC in the Browser with PyScript"&lt;/a&gt;. Published in 2022, it demonstrated something magical: running full Bayesian inference with PyMC entirely in the browser—no server, no installation, no data leaving your device. Users could define models, run NUTS sampling, and visualize posteriors, all client-side.&lt;/p&gt;
&lt;p&gt;I was excited to try it out. But when I attempted to run the examples, I discovered they no longer worked. The Python package ecosystem had evolved, dependencies had shifted, and the Pyodide environment had changed. What was once a breakthrough demo had quietly broken.&lt;/p&gt;
&lt;p&gt;So I did what any curious engineer would do on a weekend: I dove down the rabbit hole to figure out how to make it work again.&lt;/p&gt;
&lt;h2 id="the-core-challenge-getting-pytensor-to-build-for-webassembly"&gt;The core challenge: getting PyTensor to build for WebAssembly&lt;/h2&gt;&lt;p&gt;PyMC depends on PyTensor, its computational backend. PyTensor is where the heavy lifting happens: it compiles mathematical expressions into optimized code (usually C or JAX) and executes them efficiently. To run PyMC in a browser via Pyodide, I first needed to make PyTensor installable in a WebAssembly environment.&lt;/p&gt;
&lt;p&gt;This wasn't just a matter of &lt;code&gt;pip install&lt;/code&gt;. PyTensor contains C and Cython extensions that must be compiled for the target platform. For WebAssembly, that means using Emscripten and the Pyodide build tooling.&lt;/p&gt;
&lt;h2 id="the-code-changes-what-i-modified-in-pytensor"&gt;The code changes: what I modified in PyTensor&lt;/h2&gt;&lt;p&gt;Working on my fork of PyTensor (&lt;a href="https://github.com/ericmjl/pytensor"&gt;ericmjl/pytensor&lt;/a&gt;), I made targeted modifications to enable WASM builds. Here's the complete diff:&lt;/p&gt;
&lt;h3 id="change-1-making-numba-optional-on-webassembly"&gt;Change 1: making Numba optional on WebAssembly&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;File:&lt;/strong&gt; &lt;code&gt;pyproject.toml&lt;/code&gt;&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="gd"&gt;-    &amp;quot;numba&amp;gt;0.57,&amp;lt;1&amp;quot;,&lt;/span&gt;
&lt;span class="gi"&gt;+    &amp;quot;numba&amp;gt;0.57,&amp;lt;1; platform_machine != &amp;#39;wasm32&amp;#39; and sys_platform != &amp;#39;emscripten&amp;#39;&amp;quot;,&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This single line change is the critical enabler. Numba, PyTensor's JIT compiler for numerical code, is not available in WebAssembly environments. There's no way to install it—it simply doesn't exist for this platform.&lt;/p&gt;
&lt;p&gt;The fix uses PEP 508 environment markers to make Numba a conditional dependency:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;platform_machine != 'wasm32'&lt;/code&gt; excludes WASM architectures&lt;/li&gt;
&lt;li&gt;&lt;code&gt;sys_platform != 'emscripten'&lt;/code&gt; adds an extra safety check for Emscripten-based builds&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Without this change, attempting to install PyTensor in Pyodide would fail immediately with a dependency resolution error. Pyodide would try to find a Numba wheel for WASM, fail, and abort the entire installation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The tradeoff, however, is that PyTensor loses its JIT compilation capabilities on WASM.&lt;/strong&gt; Operations that would be compiled to optimized native code fall back to pure Python execution. This means slower performance, and critically, PyMC's NUTS sampler won't work.&lt;/p&gt;
&lt;h3 id="change-2-adding-pixi-development-environment-configuration"&gt;Change 2: adding Pixi development environment configuration&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;File:&lt;/strong&gt; &lt;code&gt;pyproject.toml&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;I added a complete Pixi workspace configuration to &lt;code&gt;pyproject.toml&lt;/code&gt;. This provides a reproducible development environment and includes the tooling needed to build WASM wheels:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# -----------------------------------------------------------------------------&lt;/span&gt;
&lt;span class="c1"&gt;# Pixi (pixi.prefix.dev): development environment from environment.yml&lt;/span&gt;
&lt;span class="c1"&gt;# Use: pixi install &amp;amp;&amp;amp; pixi run pytest   or   pixi shell&lt;/span&gt;
&lt;span class="c1"&gt;# -----------------------------------------------------------------------------&lt;/span&gt;
&lt;span class="k"&gt;[tool.pixi.workspace]&lt;/span&gt;
&lt;span class="n"&gt;channels&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;conda-forge&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;platforms&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;linux-64&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;osx-64&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;osx-arm64&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;win-64&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.pypi-dependencies]&lt;/span&gt;
&lt;span class="n"&gt;pytensor&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;editable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;types-setuptools&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pyodide-build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=0.29.2&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;python&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=3.11,&amp;lt;3.14&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;compilers&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;numpy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=2.0.0&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;scipy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=1,&amp;lt;2&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;filelock&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=3.15&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;etuples&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;logical-unification&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;miniKanren&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;cons&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pydeprecate&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;numba&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=0.57&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;coveralls&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;diff-cover&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;mypy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pytest&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pytest-cov&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pytest-xdist&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pytest-benchmark&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pytest-mock&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pytest-sphinx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;sphinx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=5.1.0,&amp;lt;6&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;sphinx_rtd_theme&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pygments&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pydot&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;ipython&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pymc-sphinx-theme&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;sphinx-design&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;myst-nb&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;matplotlib&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;watermark&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;ruff&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pandas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pre-commit&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;packaging&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;cython&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;graphviz&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.target.linux-64.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;mkl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;mkl-service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;libblas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*mkl&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.target.win-64.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;mkl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;mkl-service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;libblas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*mkl&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.target.osx-64.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;libblas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*accelerate&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.target.osx-arm64.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;libblas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*accelerate&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.tasks]&lt;/span&gt;
&lt;span class="n"&gt;test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;pytest&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;lint&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ruff check .&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;format&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ruff format .&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;docs&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;python -m sphinx -b html ./doc ./html&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;wheel&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;python -m build --wheel&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;sdist&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;python -m build --sdist&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;wheel-wasm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;pyodide build&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Here are the key design decisions in this configuration:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Python version pinning:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;python&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=3.11,&amp;lt;3.14&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Pyodide only supports up to Python 3.13. Without this constraint, the environment might resolve to Python 3.14+, causing the WASM build to fail with: &lt;code&gt;ValueError: Python version 3.14 is not yet supported.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PyPI dependencies for building:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;[tool.pixi.pypi-dependencies]&lt;/span&gt;
&lt;span class="n"&gt;pytensor&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;editable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;types-setuptools&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;pyodide-build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;gt;=0.29.2&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This installs PyTensor in editable mode for development, includes type stubs for mypy, and adds both &lt;code&gt;build&lt;/code&gt; (standard wheel building) and &lt;code&gt;pyodide-build&lt;/code&gt; (WASM wheel building).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Platform-specific BLAS:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;[tool.pixi.target.linux-64.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;mkl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;mkl-service&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;libblas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*mkl&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;[tool.pixi.target.osx-arm64.dependencies]&lt;/span&gt;
&lt;span class="n"&gt;libblas&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;build&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;*accelerate&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Different platforms use different BLAS implementations. Linux and Windows use Intel MKL, while macOS uses Apple's Accelerate framework. These ensure the correct linear algebra library is installed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Build task:&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;wheel-wasm&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;pyodide build&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This task runs &lt;code&gt;pyodide build&lt;/code&gt;, which compiles PyTensor for WebAssembly using Emscripten.&lt;/p&gt;
&lt;h3 id="change-3-documenting-the-wasm-build-process"&gt;Change 3: documenting the WASM build process&lt;/h3&gt;&lt;p&gt;&lt;strong&gt;File:&lt;/strong&gt; &lt;code&gt;doc/dev_start_guide.rst&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;I added documentation explaining how to build WASM wheels:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="gh"&gt;Building a WebAssembly (Pyodide) wheel&lt;/span&gt;
&lt;span class="gh"&gt;-------------------------------------&lt;/span&gt;

To build a wheel targeting WebAssembly for use with &lt;span class="s"&gt;`Pyodide &lt;/span&gt;&lt;span class="si"&gt;&amp;lt;https://pyodide.org/&amp;gt;&lt;/span&gt;&lt;span class="s"&gt;`_&lt;/span&gt; (e.g. for the browser or JupyterLite), use the Pyodide build tooling. This produces a wheel in &lt;span class="s"&gt;``dist/``&lt;/span&gt; with a name like &lt;span class="s"&gt;``*-cpXXX-cpXXX-pyodide_*_wasm32.whl``&lt;/span&gt;.

&lt;span class="gs"&gt;**One-time setup: Emscripten**&lt;/span&gt;

&lt;span class="m"&gt;1.&lt;/span&gt; Install &lt;span class="nv"&gt;`pyodide-build`&lt;/span&gt; (included in the Pixi dev env, or &lt;span class="s"&gt;``pip install pyodide-build&amp;gt;=0.29.2``&lt;/span&gt;).
&lt;span class="m"&gt;2.&lt;/span&gt; Get the Emscripten version required by your pyodide-build: &lt;span class="s"&gt;``pyodide config get emscripten_version``&lt;/span&gt;.
&lt;span class="m"&gt;3.&lt;/span&gt; Install and activate that Emscripten version using the &lt;span class="s"&gt;`Emscripten SDK (emsdk) &lt;/span&gt;&lt;span class="si"&gt;&amp;lt;https://emscripten.org/docs/getting_started/downloads.html&amp;gt;&lt;/span&gt;&lt;span class="s"&gt;`_&lt;/span&gt;:

&lt;span class="p"&gt;   ..&lt;/span&gt; &lt;span class="ow"&gt;code-block&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt; &lt;span class="k"&gt;bash&lt;/span&gt;

      git&lt;span class="w"&gt; &lt;/span&gt;clone&lt;span class="w"&gt; &lt;/span&gt;https://github.com/emscripten-core/emsdk.git
      &lt;span class="nb"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;emsdk
      ./emsdk&lt;span class="w"&gt; &lt;/span&gt;install&lt;span class="w"&gt; &lt;/span&gt;&amp;lt;version&amp;gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="c1"&gt;# use the version from step 2&lt;/span&gt;
      ./emsdk&lt;span class="w"&gt; &lt;/span&gt;activate&lt;span class="w"&gt; &lt;/span&gt;&amp;lt;version&amp;gt;
      &lt;span class="nb"&gt;source&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;emsdk_env.sh

&lt;span class="m"&gt;4.&lt;/span&gt; In any shell where you want to build the wasm wheel, ensure Emscripten is on &lt;span class="s"&gt;``PATH``&lt;/span&gt; (e.g. run &lt;span class="s"&gt;``source /path/to/emsdk/emsdk_env.sh``&lt;/span&gt;).

&lt;span class="gs"&gt;**Build the wheel**&lt;/span&gt;

From the project root, with Emscripten activated and your dev environment active (e.g. &lt;span class="s"&gt;``pixi shell``&lt;/span&gt;):

&lt;span class="p"&gt;..&lt;/span&gt; &lt;span class="ow"&gt;code-block&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt; &lt;span class="k"&gt;bash&lt;/span&gt;

   pyodide&lt;span class="w"&gt; &lt;/span&gt;build

Or with Pixi: &lt;span class="s"&gt;``pixi run wheel-wasm``&lt;/span&gt;.

The wheel will appear in &lt;span class="s"&gt;``dist/``&lt;/span&gt;. PyPI does not yet accept emscripten/wasm32 wheels; host the file elsewhere (e.g. GitHub Releases) and install in Pyodide with &lt;span class="s"&gt;``micropip.install(url)``&lt;/span&gt;. See &lt;span class="s"&gt;`Pyodide: building packages &lt;/span&gt;&lt;span class="si"&gt;&amp;lt;https://pyodide.org/en/stable/development/building-packages-from-source.html&amp;gt;&lt;/span&gt;&lt;span class="s"&gt;`_&lt;/span&gt; for details.
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This documentation walks through the Emscripten setup, the build command, and importantly, notes that PyPI doesn't accept WASM wheels yet—you need to distribute them via GitHub Releases or similar and install with &lt;code&gt;micropip.install(url)&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="what-i-actually-pr-d-to-pytensor"&gt;What I actually PR'd to PyTensor&lt;/h2&gt;&lt;p&gt;The changes above represent my weekend exploration, but they weren't what I ultimately contributed back to PyTensor. The Pixi configuration, in particular, was too large of a departure from PyTensor's existing toolchain. PyTensor uses mamba (via &lt;code&gt;environment.yml&lt;/code&gt;) for its development environment, and switching to Pixi would have been a significant change to impose on a project I don't maintain.&lt;/p&gt;
&lt;p&gt;Instead, I re-did the infrastructure changes using &lt;code&gt;pyodide-build&lt;/code&gt; while respecting PyTensor's existing mamba-based workflow. The core change (making Numba optional on WebAssembly) remained, but the development environment configuration was adapted to work with what PyTensor already had in place.&lt;/p&gt;
&lt;p&gt;This is a common lesson in open-source contribution: meeting maintainers where they are matters more than introducing your preferred tooling. The weekend experiment taught me what was needed; the PR reflected what was appropriate. You can see the actual PR here: &lt;a href="https://github.com/pymc-devs/pytensor/pull/1960"&gt;pytensor #1960&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="unfortunately-for-now-nuts-is-gone"&gt;Unfortunately (for now), NUTS is gone&lt;/h2&gt;&lt;p&gt;Unfortunately, &lt;strong&gt;NUTS (No-U-Turn Sampler) doesn't work in WASM&lt;/strong&gt;. 😭&lt;/p&gt;
&lt;p&gt;NUTS is the crown jewel of PyMC. It's the adaptive Hamiltonian Monte Carlo sampler that makes Bayesian inference efficient and robust. The 2022 PyMC Labs demo used NUTS to sample from posteriors in real-time in the browser.&lt;/p&gt;
&lt;p&gt;But here's the thing: this isn't just about Numba being unavailable. The real issue is that none of the modern MCMC sampling backends have WASM support:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;JAX&lt;/strong&gt; (used by NumPyro and BlackJAX) has an open GitHub issue &lt;a href="https://github.com/jax-ml/jax/issues/1472"&gt;#1472&lt;/a&gt; from 2019 titled "Jax for Web? (JS api or web assembly guide)" that's still open with no official WASM support&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;nutpie&lt;/strong&gt; (the Rust-based NUTS implementation) doesn't have a WASM build readily available&lt;/li&gt;
&lt;li&gt;The computational demands of Hamiltonian dynamics—computing gradients, simulating trajectories, adapting step sizes—require optimized backends that don't exist in WASM environments&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This weekend's exploration shows the path to install PyMC in WASM, but you can't use its best sampler. It's like getting a Ferrari delivered to your house, but the dealer forgot to include the keys. You can sit in it, admire the leather seats, and maybe even turn on the radio. But you're not going anywhere fast.&lt;/p&gt;
&lt;p&gt;This represents a fundamental infrastructure gap, not just a missing dependency. Getting NUTS in the browser will require either WASM ports of JAX or nutpie, or entirely new sampling backends designed for browser environments.&lt;/p&gt;
&lt;h2 id="what-does-work"&gt;What does work?&lt;/h2&gt;&lt;p&gt;Despite the NUTS heartbreak, this wasn't a failed experiment:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;PyTensor now installs in WASM environments.&lt;/strong&gt; This is non-trivial. PyTensor has C and Cython extensions that need to compile for WebAssembly. Getting that build pipeline working required understanding Pyodide's build system, setting up Emscripten correctly, and making Numba optional.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;PyMC can technically be imported.&lt;/strong&gt; Once PyTensor was installable, PyMC followed. You can define models, create random variables, and work with the API. The foundation is there.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alternative samplers might still work.&lt;/strong&gt; While NUTS is off the table, other samplers—like Metropolis-Hastings or Slice sampling—might be viable for small models. They're slower and less robust than NUTS, but they don't require JIT compilation. I didn't test, but I think this will hold true!&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The roadmap is clearer.&lt;/strong&gt; If someone wants to bring full PyMC to the browser, the path forward is documented. It requires either (a) building WASM support into JAX (a massive undertaking that's been an open request since 2019), (b) creating WASM builds for nutpie, or (c) building entirely new sampling backends designed for browser environments (also non-trivial, but potentially more feasible).&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="a-weekend-well-spent"&gt;A weekend well spent&lt;/h2&gt;&lt;p&gt;Did I achieve my original goal of running PyMC in the browser with NUTS sampling? No. The technical limitations of WASM environments made that impossible with the current architecture.&lt;/p&gt;
&lt;p&gt;But that's the nature of weekend experiments. You explore, you hit walls, you learn. I now understand PyTensor's dependency structure at a deeper level. I've learned how Pyodide builds work and the constraints they impose. I've identified the broader infrastructure gap (MCMC sampling backends lacking WASM support) that needs solving for true browser-based Bayesian inference.&lt;/p&gt;
&lt;p&gt;The dream of running PyMC entirely in the browser isn't dead—it's just waiting for the right infrastructure. Until JAX or nutpie (or something else) supports WASM, we'll keep pushing that car downhill.&lt;/p&gt;
</content></entry><entry><title>Mastering Personal Knowledge Management with Obsidian and AI</title><link href="https://ericmjl.github.io/blog/2026/3/6/mastering-personal-knowledge-management-with-obsidian-and-ai/" rel="alternate"/><updated>2026-03-06T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:0a70d5db-2590-3622-84d3-785f1bf45d29</id><content type="html">&lt;p&gt;Folks have asked me how I do personal knowledge management (PKM) at work. The question becomes more pressing when they learn how many projects and people I need to interact with on a weekly basis. At the time of writing, I manage twelve people across two teams, each handling 2-4 projects of their own. That's a lot of context to keep straight.&lt;/p&gt;
&lt;p&gt;I decided to document what I'm doing for PKM. Hopefully it serves as inspiration for you.&lt;/p&gt;
&lt;p&gt;I've written before about &lt;a href="../../../../2020/12/15/building-a-personal-knowledge-graph-on-obsidian/"&gt;why I chose Obsidian&lt;/a&gt;; this post shows how that decision evolved with AI integration over five years.&lt;/p&gt;
&lt;h2 id="the-plain-text-decision"&gt;The plain text decision&lt;/h2&gt;&lt;p&gt;In 2022, I decided to make personal knowledge management a priority at work. I faced a choice: Confluence, OneNote, or a new kid on the block, Obsidian. I chose plain text and graphs. I chose Obsidian. I chose not to lock my data inside a vendor system. I chose freedom and sovereignty for my information.&lt;/p&gt;
&lt;p&gt;That decision was prescient in ways I couldn't have predicted. Most of us normies back then wouldn't have guessed that plain text would be exactly the right format for 2025 and 2026 era knowledge management. The visionaries saw it coming; I just got lucky because I loved the graph view in Obsidian and thought of it as a really cool tool. But holy smokes, has that choice paid off.&lt;/p&gt;
&lt;p&gt;Text files are as primitive as it gets: no proprietary formats, no vendor lock-in, just files that can be read on any system. When AI coding agents arrived, my vault was already in a format they could process natively. No migration needed. No conversion layer. No API integration. The simplicity I chose became an unlock I never planned for.&lt;/p&gt;
&lt;h2 id="the-core-system"&gt;The core system&lt;/h2&gt;&lt;p&gt;My Obsidian vault is built around distinct note types. Monthly collections of daily bullet journals capture my day-to-day activities, one note per month with a running log of meetings and work. Meeting notes follow a structured template. People notes are dossiers for everyone I work with (to put it in CIA terms, I keep a file on everyone I interact with regularly). Project notes act as control towers, linking out to meetings, people, and status updates. A miscellaneous collection handles everything else. The structure was inspired in part by Thiago Forte's numbered folder system, though I've simplified it over time.&lt;/p&gt;
&lt;p&gt;The most important thing isn't my specific implementation. It's that I have a system at all, and it's documented in an &lt;code&gt;AGENTS.md&lt;/code&gt; file so my coding agents understand it too.&lt;/p&gt;
&lt;h2 id="ingesting-information"&gt;Ingesting information&lt;/h2&gt;&lt;p&gt;The lifecycle of my workflow starts with ingestion. Meeting notes arrive as transcripts or AI-generated summaries. In the past, structuring these was tedious work. Now I paste them into OpenCode and my meeting notes skill handles the rest.&lt;/p&gt;
&lt;p&gt;The skill knows the template I want. It handles various input formats: AI-generated summaries, transcripts with good speaker assignments, and transcripts with poor speaker assignments. I flag the quality when I know it's bad. The skill extracts key information and formats everything consistently. For one-on-ones, it ensures notes are attached to both the meeting log and the person's individual page, so I can track the full history of our conversations.&lt;/p&gt;
&lt;p&gt;Beyond meetings, I ingest PowerPoints, Word docs, PDFs, and Excel spreadsheets into my vault as contextual information. The key insight is getting everything into plain text format. For Word documents, a Python script converts them to plain text using &lt;code&gt;python-docx&lt;/code&gt;, which is then printed to the terminal or dumped to disk at &lt;code&gt;/tmp&lt;/code&gt;, both of which are readable by a coding agent. Even lightly misformatted plain text contains enough information density for a coding agent to read and summarize.&lt;/p&gt;
&lt;p&gt;For PowerPoints, I use dual parsing paths. One path extracts the XML structure directly using &lt;code&gt;python-pptx&lt;/code&gt;. The second path converts each slide to an image using &lt;code&gt;libreoffice&lt;/code&gt; and &lt;code&gt;PIL&lt;/code&gt;, captions it with a vision-language model via APIs, and strings the captions together into a coherent narrative. Combined, I estimate that I can get a 90-95% accurate textual representation. PDFs follow a similar pattern: text extraction for normal PDFs, image captioning for scanned documents.&lt;/p&gt;
&lt;p&gt;Excel spreadsheets are read directly by the coding agent using &lt;code&gt;openpyxl&lt;/code&gt;, not &lt;code&gt;pandas&lt;/code&gt;. The key difference matters: &lt;code&gt;pandas&lt;/code&gt; assumes an established table structure, but real-world spreadsheets are messy. With &lt;code&gt;openpyxl&lt;/code&gt;, the agent can read the granular cellular structure across each sheet, identifying merged cells, free text scattered in random locations, and arbitrary layouts. This structural mapping follows a progressive reveal principle: the agent first identifies the spreadsheet's architecture without necessarily reading every cell's contents, then zooms into relevant sections. This approach handles the chaos of actual spreadsheets far better than forcing everything into a tabular assumption. It's powerful when I need to understand financial data without being a finance person.&lt;/p&gt;
&lt;h2 id="managing-and-maintaining"&gt;Managing and maintaining&lt;/h2&gt;&lt;p&gt;With information in the vault, the next phase is keeping it current. With twelve people across two teams, there are a lot of details I don't pick up or retain in my working memory. That's why external memory matters. Without it, things would fall through the cracks.&lt;/p&gt;
&lt;p&gt;When I hit a context block (when I look up a project or person and realize something's missing), I trigger a "sweep". My instructions to the coding agent are to update my people notes and/or project notes based on source material present in the vault. People and project notes are always derivative from sources, so any updates must include quotations from those source notes. I stay in the loop for verification. Hallucinations are rare, maybe once every four or five sweeps, and usually trace back to inaccurate transcripts rather than agent errors.&lt;/p&gt;
&lt;p&gt;This is incredibly helpful for how I interact with people. My assumption is that I'm going to be forgetful. My external memory will be approximately correct, and I have a process for keeping it refined over time. So I can rely more on the vault instead of second-guessing myself based on incomplete memory. It tempers how I think about interacting with someone, not by changing my mind about them, but by giving me confidence that I'm not missing something important.&lt;/p&gt;
&lt;p&gt;There are ethical boundaries. I don't capture personal details if people aren't comfortable with that. The dossiers are professional, not invasive.&lt;/p&gt;
&lt;p&gt;Periodically, I do retrieval practice. This is how we make information stick; read "Make It Stick" to learn more. Review looks like this: I take my people notes and project notes and ask what's missing. Is there a piece of knowledge I remember that isn't captured? If yes, I fill in the blanks. I also check whether claims are substantiated with links and quotes. This fact-checking pass keeps the vault trustworthy and protects me from remembering something erroneous. A spell-checker list handles transcription errors, and my &lt;code&gt;AGENTS.md&lt;/code&gt; links to &lt;code&gt;HEARTBEAT.md&lt;/code&gt; to sanitize the vault of inaccurate information.&lt;/p&gt;
&lt;h2 id="producing-and-sharing"&gt;Producing and sharing&lt;/h2&gt;&lt;p&gt;The final phase is producing outputs for others. I curate what gets published rather than exporting everything. The agent creates a publishable version based on my guidance. I haven't settled on hard rules for curation yet. I'd rather review and decide at publish time than tag things as publishable during capture. That workflow feels right to me.&lt;/p&gt;
&lt;p&gt;For Confluence, a Python script publishes markdown directly, with YAML front matter defining the space and parent page. For GitHub users, notes can become Gists via the GitHub CLI. With the appropriate skills, Markdown files transform into HTML presentations, and with web technologies, those presentations become interactive. For Jira, a colleague created a skill that writes Jira tickets. We firmly believe that humans shouldn't be filling forms out; AI should be filling forms for us.&lt;/p&gt;
&lt;p&gt;PowerPoint decks can be generated via Python scripts. Word documents come from markdown via Pandoc. The scripts run with uv, and LibreOffice handles conversions.&lt;/p&gt;
&lt;p&gt;Each script maintains its own environment using PEP 723 inline script metadata. This means dependencies are declared at the top of each script in a special comment block:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# /// script&lt;/span&gt;
&lt;span class="c1"&gt;# dependencies = [&amp;quot;python-docx&amp;quot;, &amp;quot;python-pptx&amp;quot;, &amp;quot;pandas&amp;quot;]&lt;/span&gt;
&lt;span class="c1"&gt;# ///&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;When I run &lt;code&gt;uv run script.py&lt;/code&gt;, uv automatically creates an isolated environment with just those dependencies, executes the script, and cleans up. No virtual environments to manage. No &lt;code&gt;requirements.txt&lt;/code&gt; files scattered everywhere. No "works on my machine" problems.&lt;/p&gt;
&lt;h2 id="the-role-of-agent-skills"&gt;The role of agent skills&lt;/h2&gt;&lt;p&gt;Agent skills effectively encode procedural knowledge into executable markdown. Over time, it compounds; on fewer and fewer occasions do I need to repeat instructions, which is incredibly liberating. The model infers which skill to use most of the time. When it doesn't, I correct it explicitly and ask the coding agent to update the skill file for the future as well.&lt;/p&gt;
&lt;p&gt;Designing a skill means thinking about the desired output and the tools needed to get there. I discover edge cases in the wild and update immediately. The earlier errors are caught, the better.&lt;/p&gt;
&lt;h2 id="what-s-still-friction"&gt;What's still friction&lt;/h2&gt;&lt;p&gt;One pain point remains. I want to ingest Office files by pasting a URL, but I still need to download a copy first, then feed that copy to the agent skill. Programmatic access to cloud documents would eliminate this step. From the user side, nothing else would change. I'd just paste the URL and go.&lt;/p&gt;
&lt;p&gt;But even with this friction, the system pays for itself. Knowledge management overhead dropped from thirty to forty percent of my time down to less than ten percent. I fix errors as I encounter them rather than scheduling dedicated maintenance. That recovered bandwidth goes toward better thinking and context gathering.&lt;/p&gt;
&lt;h2 id="getting-started"&gt;Getting started&lt;/h2&gt;&lt;p&gt;What stops people from building systems like this? I believe it is two things: imagination and technical skill. You need imagination to envision converting diverse file formats into plain text. You need technical skill to know that it's possible.&lt;/p&gt;
&lt;p&gt;The two feed each other. I experienced this with web technologies. Before I got familiar with building stuff on the web, I wondered what was even possible. Once I actually built things, I knew. Technical skill feeds your imagination, and imagination drives you to learn more technical skills.&lt;/p&gt;
&lt;p&gt;For those starting without technical skills, use AI to learn programming. Find a language with a supportive human community to verify what you learn. AI hallucinates, and you need other people around you to help apply judgment and skill to AI outputs. You also need critical thinking skills and the initiative to act on what agents produce.&lt;/p&gt;
&lt;h2 id="skills-you-can-use-today"&gt;Skills you can use today&lt;/h2&gt;&lt;p&gt;If you want to experiment with agent skills, here are some I've published:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/ericmjl/skills/tree/main/skills/html-presentations"&gt;html-presentations&lt;/a&gt; - Turn markdown into gorgeous HTML slides&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ericmjl/skills/tree/main/skills/gh-daily-timeline"&gt;gh-daily-timeline&lt;/a&gt; - See your GitHub activity for any given day&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ericmjl/skills/tree/main/skills/gh-activity-summary"&gt;gh-activity-summary&lt;/a&gt; - Generate a plain-language summary of your GitHub work over any time period&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ericmjl/skills/tree/main/skills/publish-to-google-docs"&gt;publish-to-google-docs&lt;/a&gt; - Push markdown notes to Google Docs&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-bigger-picture"&gt;The bigger picture&lt;/h2&gt;&lt;p&gt;With such a system in place, repetitive, monotonous, and manual work can be offloaded to computers and AI. With a personal knowledge system, we can carry a broader scope of responsibilities and grow into new challenges for two reasons: we can externalize our memory more easily, and we can format information in ways that fit our brains.&lt;/p&gt;
&lt;p&gt;I'm not asking people to do more at the same time. I'm asking them to expand their dynamic range over time so they're not stuck doing the same old boring thing over and over. That repetitive monotonous stuff should have been given away to AI and computers a long time ago.&lt;/p&gt;
&lt;p&gt;This is useful for your career. It keeps things interesting. Every day that you make an incremental but permanent improvement compounds over time.&lt;/p&gt;
&lt;p&gt;The vignettes I've shared are not a prescription. Rather, I hope you treat them as an invitation. Plain text plus coding agents is a powerful combination. Your system will look different from mine, and that's part of the point. Experiment and explore, and find what works for you.&lt;/p&gt;
</content></entry><entry><title>How to stay in control when doing EDA with coding agents</title><link href="https://ericmjl.github.io/blog/2026/2/13/agentic-eda/" rel="alternate"/><updated>2026-02-13T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:f7d7fdf8-57b8-3130-99d4-c72d42aea0cc</id><content type="html">&lt;p&gt;Speed without control is just chaos.&lt;/p&gt;
&lt;p&gt;I've seen teammates compress a week and a half of analysis work into half a day using coding agents. That's a 5-10x speedup. But here's the thing: that speed only matters if you stay in the driver's seat. Otherwise you're not doing data science, you're just generating artifacts.&lt;/p&gt;
&lt;p&gt;The real unlock isn't that agents write code fast. It's that they can be guided through a structure that keeps you in control of the analysis.&lt;/p&gt;
&lt;h2 id="the-problem-isn-t-speed-it-s-agency"&gt;The problem isn't speed, it's agency&lt;/h2&gt;&lt;p&gt;Coding agents are eager. Give them a CSV file and they'll open it, generate a dozen plots, and dump a wall of code before you've finished describing what you're actually looking for. That feels productive. It isn't.&lt;/p&gt;
&lt;p&gt;The problem is that you've lost the thread. You didn't formulate a clear question. You didn't think through what the x-axis and y-axis should be. You're now reacting to whatever the agent produced, rather than steering toward an answer.&lt;/p&gt;
&lt;p&gt;I've developed a different approach, codified in two skills I use with my coding agents: &lt;a href="https://github.com/ericmjl/skills/tree/main/skills/scientific-eda"&gt;scientific-eda&lt;/a&gt; for exploratory data analysis and &lt;a href="https://github.com/ericmjl/skills/tree/main/skills/ml-experimentation"&gt;ml-experimentation&lt;/a&gt; for machine learning experiments. The pattern is the same in both: slow down first, gate on artifacts (plots, tables, etc.), and structure the session so both you and the agent can follow what happened.&lt;/p&gt;
&lt;h2 id="slow-down-first-the-socratic-opening"&gt;Slow down first: the Socratic opening&lt;/h2&gt;&lt;p&gt;The first design principle is counterintuitive: slow down before you speed up.&lt;/p&gt;
&lt;p&gt;When you invoke the scientific-eda skill, the agent does not immediately load your data and start plotting. Instead, it asks you questions. What's the problem context? What are you hoping to learn or decide? What constraints matter?&lt;/p&gt;
&lt;p&gt;From the skill definition:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;Do not open the data file and start coding or plotting. Ask for or confirm: the problem context—biological, chemical, or data-science question; what the user hopes to learn or decide; and any constraints.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There's also an explicit guardrail: "ask 'why' before executing." When you request a specific plot or table, the agent briefly asks what question or decision it serves. This isn't bureaucracy. It's alignment. The agent is checking that you've thought through the request before it spends your time executing it.&lt;/p&gt;
&lt;p&gt;This Socratic opening feels slower. But it prevents the far more common waste: generating plots you didn't need, going down rabbit holes you can't explain, and ending up with a folder full of artifacts and no clear answer.&lt;/p&gt;
&lt;h2 id="gate-everything-on-artifacts"&gt;Gate everything on artifacts&lt;/h2&gt;&lt;p&gt;The second design principle is more specific: one artifact at a time.&lt;/p&gt;
&lt;p&gt;If you can't describe what you want, you're not ready to execute. The agent waits. You think. You describe. Then the agent generates exactly what you asked for.&lt;/p&gt;
&lt;p&gt;For a plot, this means articulating the x-axis, the y-axis, and what pattern you're looking for. What would confirm or refute your hypothesis? If you can't answer, the analysis isn't ready to run.&lt;/p&gt;
&lt;p&gt;For a table, this means specifying the columns, rows, and aggregation level. What comparison are you trying to make? What decision will this table inform? A vague request like "show me the data" isn't actionable. "Show me the mean expression level by treatment group" is.&lt;/p&gt;
&lt;p&gt;This is a forcing function for clarity. Describing an artifact precisely forces you to articulate the question you're actually asking. The tradeoff is worth it. You give up a bit of speed up front for precision in execution. And because the agent can generate code in seconds rather than minutes, the net result is still a massive speedup. My teammates went from 1.5-2 weeks to half a day. The precision tax is negligible compared to the execution dividend.&lt;/p&gt;
&lt;h2 id="the-session-structure"&gt;The session structure&lt;/h2&gt;&lt;p&gt;Here's where structure becomes a feature, not overhead.&lt;/p&gt;
&lt;p&gt;Each analysis session is a timestamped folder. The naming convention is ISO datetime plus a descriptive slug:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;analysis/
  2025-02-05T14-30-00-protein-binding/
    journal.md      # append-only; shape, actions, findings
    plots/          # WebP figures only
    scripts/        # disposable PEP723 scripts; uv run
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Let's walk through what each piece does.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;journal.md&lt;/strong&gt; is the memory. Before each action, the agent reads the journal. After each action, it appends what happened. Entries get timestamped and tagged: &lt;code&gt;[SHAPE]&lt;/code&gt; for data structure discoveries, &lt;code&gt;[PLOT]&lt;/code&gt; for visualizations, &lt;code&gt;[FINDING]&lt;/code&gt; for observations, &lt;code&gt;[NEXT]&lt;/code&gt; for suggested next steps. The journal is scannable. It's also the entry point for anyone (including future you) who wants to understand what happened without reading the code.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;plots/&lt;/strong&gt; holds all figures from the session. The skill specifies WebP format for smaller file sizes, though that's a minor detail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;scripts/&lt;/strong&gt; contains disposable Python scripts. Each script has PEP723 inline metadata at the top, declaring its own dependencies. You run them with &lt;code&gt;uv run script.py&lt;/code&gt; from the session folder. No environment wrangling. No "which virtualenv am I in?" confusion. One script, one plot.&lt;/p&gt;
&lt;p&gt;The structure serves two purposes. First, it gives the agent a clear protocol to follow: read journal, execute, append to journal, suggest next step. Second, it leaves a trace that any human can follow. You can read the plan, then the journal, then the report, and understand the entire analysis without touching the code.&lt;/p&gt;
&lt;h2 id="what-changes-in-team-conversations"&gt;What changes in team conversations&lt;/h2&gt;&lt;p&gt;Something unexpected happened when I started using this approach with teammates.&lt;/p&gt;
&lt;p&gt;The conversations changed. We stopped asking "why did you write it that way?" and started asking "why did the agent write it that way?" The ego attached to code ownership evaporated. We could critique the work without critiquing each other.&lt;/p&gt;
&lt;p&gt;This matters more than I expected. In the pre-agent world, I'd invest 50-70% of my mental energy on implementation details: wrangling data frames, handling edge cases, debugging syntax errors. That labor created attachment. When someone questioned my code, it felt like they were questioning my thinking.&lt;/p&gt;
&lt;p&gt;Now the agent writes the code. I focus on the questions and verify the work. My teammates and I have more productive scientific conversations because we're discussing the analysis, not defending the implementation. We check the agent's work together, and if something's wrong, we just ask the agent to fix it.&lt;/p&gt;
&lt;h2 id="design-for-the-human"&gt;Design for the human&lt;/h2&gt;&lt;p&gt;The pattern that emerged is simple: if you design for the human, the agent follows.&lt;/p&gt;
&lt;p&gt;The structure that makes your analysis traceable is the same structure that keeps the agent aligned. The journal that helps future-you understand what happened also helps the agent decide what to do next. The artifact-gating that forces you to think clearly also gives the agent precise instructions to execute.&lt;/p&gt;
&lt;p&gt;You stay in control by slowing down at the decision points, describing what you want before you get it, and keeping a running record of what happened. The agent becomes a force multiplier rather than a loose cannon.&lt;/p&gt;
&lt;p&gt;The skills I've linked here are just text files. They're prompts, structured in a way that an LLM-based coding agent can follow. You can copy them, modify them, or write your own. The key insight isn't in any particular skill, it's in the design pattern: gate analysis on artifacts, structure sessions with journals, and make the human's job explicit before the agent's job begins.&lt;/p&gt;
&lt;p&gt;The 5-10x speedup is real. But the real win is that you get to stay the scientist.&lt;/p&gt;
</content></entry><entry><title>How to Do Agentic Data Science</title><link href="https://ericmjl.github.io/blog/2026/2/1/how-to-do-agentic-data-science/" rel="alternate"/><updated>2026-02-01T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:60e44946-3c79-3457-bbcf-b10ead198a9e</id><content type="html">&lt;p&gt;Having tasted what agentic coding could look like for software development, I wanted to know what it would look like for data science - this meant training machine learning models and answering scientific questions. So I started experimenting, at work, and on my own at home as well. Here are ten lessons I've learned from my experiments thus far.&lt;/p&gt;
&lt;h2 id="1-be-prescriptive-in-your-prompting"&gt;1. Be prescriptive in your prompting&lt;/h2&gt;&lt;p&gt;Similar to building software, you need to know exactly what you want and how you'll evaluate the outcome. The difference, however, is as follows: With software, you will often know what you need to build, but with data science, you can only know what hypotheses need to be verified, which means you will need to iterate your way to the answer. Nonetheless, it is possible to leverage coding agents to move quickly.&lt;/p&gt;
&lt;p&gt;The parallels are striking: if you frame each question you ask in terms of an observable outcome, you can set up your coding agent to write code that produces an output that can be evaluated for correctness, just like with software tests!&lt;/p&gt;
&lt;p&gt;Here, your ability to describe precisely the hypothesis you're exploring, and the ability to describe in precise language what the answer would look like if the hypothesis held true or not, are critical components of what enables the coding agent to figure out what needs to be counterfactually true (within the codebase or the data) in order for your hypothesis to hold true.&lt;/p&gt;
&lt;p&gt;Here is an example from my work. In a machine learning experiment with synthetic data, I wanted to hit 100% sequence editing performance. (It was synthetic data after all!) The coding agent hit a scenario where it was only doing 25%. With the clear goal in mind, it proposed edits to the code, edited the code, and re-ran experiments until it hit 100%. All without cheating; I know, because I checked!&lt;/p&gt;
&lt;h2 id="2-strong-patterns-in-the-file-system"&gt;2. Strong patterns in the file system&lt;/h2&gt;&lt;p&gt;The agent, like humans, needs a predictable place for experiments. Similar to how a software repo has a conventional layout (src/, tests/, and so on), your experiments need a conventional layout so the agent knows where to put things and where to look. Within the &lt;a href="https://github.com/ericmjl/skills/tree/main/skills/ml-experimentation"&gt;experimentation skill that I wrote&lt;/a&gt;, I instruct the coding agent to do its work inside an &lt;code&gt;experiments&lt;/code&gt; folder. Underneath that, for each experiment, we have datetime-prefixed subfolders, in which, there's a README file, a &lt;code&gt;plots&lt;/code&gt; directory, a &lt;code&gt;data&lt;/code&gt; directory, a &lt;code&gt;scripts&lt;/code&gt; directory. Naming things logically helps, but the scheme matters more than the exact names. Coding agents will follow the patterns you already have.&lt;/p&gt;
&lt;h2 id="3-put-logging-instructions-in-agents.md"&gt;3. Put logging instructions in AGENTS.md&lt;/h2&gt;&lt;p&gt;With software, the feature one ask's a coding agent to build either succeeds or fails, and this can be automatically verified using programmatically-runnable unit and integration tests. With data science, experimental runs produce logs and metrics, but aren't easily boolean pass/fail like software tests. In both cases, however, and the agent can introspect logs to figure out what to change!&lt;/p&gt;
&lt;p&gt;Your AGENTS.md file should include instructions for putting enough logging in place so the LLM can introspect what's going on during the experiment. I've written elsewhere about &lt;a href="../../../../2025/10/4/how-to-teach-your-coding-agent-with-agentsmd/"&gt;how to teach your coding agent with AGENTS.md&lt;/a&gt; and &lt;a href="../../../1/17/how-to-build-self-improving-coding-agents-part-1/"&gt;using AGENTS.md as repository memory&lt;/a&gt; for self-improving agents. Pair that with tools that run code in the terminal so the agent gets logs it can read. When the agent can read the logs, it can figure out what's wrong and what to change.&lt;/p&gt;
&lt;p&gt;In my work, logging and printing to terminal were what let my agent fix a masking strategy that was only yielding 25% correctness. It read the logs, proposed a fix, re-ran, and got to where we needed to be. No intervention on my part. A 3 day experiment became 20 minutes.&lt;/p&gt;
&lt;h2 id="4-give-it-report-writing-skills"&gt;4. Give it report-writing skills&lt;/h2&gt;&lt;p&gt;The agent can write code and read mountains of logs, but you need something else: a human-readable summary of what it observed and what looked weird, so you can triage without re-reading every log. Give your coding agent instructions (e.g. in an &lt;a href="https://agentskills.io/home"&gt;agent skill&lt;/a&gt;) to write out in plain language what it observed during the model evaluation phase. It should read execution logs throughout the run. Tell it to write down anything that looks weird for follow-up. If something is off, it should say so. You get a readable summary and a list of things to dig into.&lt;/p&gt;
&lt;p&gt;For reports (e.g. &lt;code&gt;reports.md&lt;/code&gt;), encode in the skill that every table and every plot must be scrutinized. Ensure that plots are generated for every table, and that someone (you or the agent, with you verifying) carefully checks for inconsistencies between AI-generated plots and the tables they are supposed to reflect. The agent can miss things. It is valid to ask the AI to check its own work, but only if you have an idea of exactly where it is wrong and you tell it as such. Vague "double-check this" rarely helps; "the values in figure 2 do not match the second column of table 1" gives the agent something it can fix.&lt;/p&gt;
&lt;h2 id="5-have-the-agent-keep-an-append-only-journal-of-observations"&gt;5. Have the agent keep an append-only journal of observations&lt;/h2&gt;&lt;p&gt;Within the skill, instruct the agent to keep a single file (e.g. &lt;code&gt;notes.md&lt;/code&gt; or &lt;code&gt;journal.md&lt;/code&gt;) that it is told to only append to, never overwrite. The journal is not just for the agent. You should add to it too: things you noticed while looking at the data, gut feelings, weird patterns. It becomes a running log of what was going on, from both sides, that you can go back and summarize later. The point is to capture the thought process while you are doing the work.&lt;/p&gt;
&lt;h2 id="6-have-the-agent-generate-diagnostic-plots-for-you"&gt;6. Have the agent generate diagnostic plots for you&lt;/h2&gt;&lt;p&gt;Logs and plots are complementary: logs are agent-accessible, but plots are human-accessible versions of the same underlying performance data. Have the agent generate diagnostic plots for you. The agent can propose fixes from the logs, but it can't build your intuition; you're the one who has to smell when something is off. Nothing beats looking at the data yourself, otherwise you never build intuition for what's happening! I still looked at the logs and plots myself to make sure the metrics were real and the agent wasn't hallucinating. Your prior experience is what lets you smell when something is off.&lt;/p&gt;
&lt;h2 id="7-instruct-the-agent-to-write-the-minimalist-version-first"&gt;7. Instruct the agent to write the minimalist version first&lt;/h2&gt;&lt;p&gt;With software, you run tests in seconds. With ML, you're tempted to train for hours and rush to the real data and the real training run. As a human, you don't want to "waste" time proving out the pipeline when you could just run the full thing, but that mentality is exactly what makes you unable to debug machine learning code on a tight loop. That temptation is exactly why you should instruct the agent to do the opposite: write the minimalist version first, then use it to work out elementary errors before scaling up.&lt;/p&gt;
&lt;p&gt;That means train for one iteration, not even one epoch. Use miniature versions of the final model (e.g. a tiny custom deep net with the same architecture but a fraction of the parameters). Check for shape errors, data-loading bugs, and that the forward pass runs end to end. All the sanity checks you would do manually to prove that things work, but that you are tempted to skip. Encode in AGENTS.md that the agent must implement and run this minimal version before moving to full-scale training. The agent does not have your impatience; use that to your advantage. Once the minimal run passes, you can scale up with confidence.&lt;/p&gt;
&lt;h2 id="8-ask-the-coding-agent-to-guide-you-through-step-by-step"&gt;8. Ask the coding agent to guide you through step by step&lt;/h2&gt;&lt;p&gt;I do this often with large software refactors within &lt;code&gt;canvas-chat&lt;/code&gt;, in which I ask the coding agent to prioritize for me a list of manual checks I need to look at. This is particularly helpful when I'm (a) context switching back into the project, or (b) running on fumes but my gut tells me we're so close to the end. (Though really, you shouldn't be doing any work if you're close to dozing off at 10:30 pm...)&lt;/p&gt;
&lt;p&gt;The same applies to data science and data exploration! After having the coding agent autonomously execute on your experiment, you can have it walk through what it's done step by step, giving you the space to operate at your pace -- at the speed of your thought! Of course, if you're in a better state than merely "running on fumes", you can (and should) treat the coding agent as a research partner and ask questions back to critically evaluate whether the output is correct or not. What I have found is that there will still be unexplored paths that need to be trodden, and you can send a coding agent off on that direction on the side.&lt;/p&gt;
&lt;h2 id="9-learn-the-vocabulary-the-coding-agent-uses"&gt;9. Learn the vocabulary the coding agent uses&lt;/h2&gt;&lt;p&gt;Pay attention to the terms the agent uses when it describes what it did. You can reuse that vocabulary in future prompts and get more precise results. For example, in developing canvas-chat, I used this non-optimal verbiage in my prompt:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;ok, I see, the default node size made it such that the next/prev buttons were hidden away. Can we make the pagination controls visible regardless of the size of the node?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Cursor's agent replied with something like "making the pagination toolbar sticky". That gave me a more compact way to express exactly what I need the next time. If you don't know the vocab at first, this is a great way to expand your technical vocabulary too.&lt;/p&gt;
&lt;h2 id="10-for-exploration-treat-the-agent-as-an-executor-that-follows-your-curiosity"&gt;10. For exploration, treat the agent as an executor that follows your curiosity&lt;/h2&gt;&lt;p&gt;What you don't want in agentic data exploration is for the coding agent to hand you a boatload of output and leave you no room to follow your own curiosity. Flip the table: treat the agent as an executor of your ideas. You lead; it follows. Instruct it that it is not allowed to race ahead. It should only execute on the one thing you want, and it should ask you questions to clarify and narrow down what you actually want before it goes and does it. In other words, it is there to be a jazz partner for your data exploration.&lt;/p&gt;
&lt;p&gt;You can run that partnership a few ways. One is to have the agent write scripts that produce plots on disk; you run them, look at the output, then ask for the next thing. Another is to go one level higher and work inside &lt;a href="https://marimo.io"&gt;Marimo&lt;/a&gt; notebooks, using Marimo's reactive execution so you go one cell at a time, one question at a time. I've written about &lt;a href="../../../../2025/10/28/use-coding-agents-to-write-marimo-notebooks/"&gt;using coding agents to write Marimo notebooks&lt;/a&gt; if you want to try that path.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The agents handle the implementation. You handle the inquiry. The ten practices above - prescriptive goals, clear structure, logging, reports, an append-only journal, diagnostic plots and your own eyes on the data, the minimalist version first, having the agent guide you step by step when you need it, learning the agent's vocabulary, and in exploration keeping the agent as your jazz partner - are what make that partnership work. I've spent nearly a decade training ML models by hand, so I know what I want, and I have developed a sense of taste for what success looks like. You can get to the same level of taste with AI assistance, but you must work for it. I'll write separately about how I'm learning new things with AI. The point is not to hand off the science, but to do more of it!&lt;/p&gt;
</content></entry><entry><title>Model feel, fast tests, and AI coding that stays in flow</title><link href="https://ericmjl.github.io/blog/2026/1/25/model-feel-fast-tests-and-ai-coding-that-stays-in-flow/" rel="alternate"/><updated>2026-01-25T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:e17e872c-bf5f-3a8b-8ad9-789f8afa713d</id><content type="html">&lt;p&gt;Most of the conversation about AI coding models focuses on performance metrics. Benchmarks, evals, pass rates, latency. Useful stuff, but it misses the part that actually shapes my day-to-day: what it &lt;em&gt;feels&lt;/em&gt; like to work with the model.&lt;/p&gt;
&lt;p&gt;Once you start using LLMs as coding agents, the qualitative experience becomes a throughput issue. It affects how often you intervene, how much you trust what is happening, and whether you stay in flow or spend your time cleaning up weird breakage.&lt;/p&gt;
&lt;p&gt;Two axes keep showing up for me.&lt;/p&gt;
&lt;p&gt;First is time horizon and supervision style: long-horizon autonomy versus short-horizon iteration.&lt;/p&gt;
&lt;p&gt;Second is personality and verbosity: how the model behaves when it is wrong, how much it narrates, and whether it stays constructive or spirals into apology loops.&lt;/p&gt;
&lt;p&gt;There is also a third ingredient that ends up mattering as much as the model: the agentic harness. By that I mean the tools and checks that the agent can run to verify it did not break behavior, &lt;em&gt;and&lt;/em&gt; whether the harness gives you streaming and visual feedback—a live trace of what the model is doing—or leaves you staring at a spinner until the answer drops. Good harness beats model swapping more often than I expected.&lt;/p&gt;
&lt;h2 id="long-horizon-autonomy-vs-short-horizon-iteration"&gt;Long-horizon autonomy vs short-horizon iteration&lt;/h2&gt;&lt;p&gt;I call it "Opus-feel" when a model has that "ask and it shall be given" vibe with a longer time horizon. You describe what you want, it runs for a while, and it comes back with a plausible scaffold. It is great for momentum.&lt;/p&gt;
&lt;p&gt;I call it "Sonnet-feel" when a model leans toward shorter-horizon iteration. It works better when you are walking through a real codebase step by step, keeping changes small enough that you can validate what happened, correct course, and keep going.&lt;/p&gt;
&lt;p&gt;Another way to put it is that long-horizon autonomy pushes you toward a spec-and-review loop, while short-horizon iteration pushes you toward a steer-and-verify loop. Both can be productive. They just fail differently.&lt;/p&gt;
&lt;p&gt;In a sufficiently large codebase, you cannot rely solely on long-horizon autonomy where you ask for something with a vague description and hope it lands cleanly. You are not always guaranteed something well organized, especially when the job is refactoring rather than greenfield scaffolding.&lt;/p&gt;
&lt;p&gt;A concrete example for me came from Canvas chat. At the time, everything was tied to &lt;code&gt;app.js&lt;/code&gt; and &lt;code&gt;app.py&lt;/code&gt;. When I wanted to refactor things into plugins, I needed to dogfood a plugin pattern in the codebase itself.&lt;/p&gt;
&lt;p&gt;Long-horizon autonomy struggled here. It could generate a plugin pattern, but it was not great at the careful, incremental work of extracting behavior out of a monolith and into a clean plugin boundary.&lt;/p&gt;
&lt;p&gt;Walking bit by bit with Sonnet or Sonnet-quality models was a very different experience. The big win was that I could study the LLM traces live (the tool calls and file edits it proposes step by step) and see where edits were being made. If I noticed a feature handler getting added to &lt;code&gt;app.js&lt;/code&gt; when it clearly belonged in a plugin file, I could intervene immediately and ask, "Why is that thing over there in &lt;code&gt;app.js&lt;/code&gt;? Why is it not inside the plugin file instead?" That kind of interactive, traceable work is where the short-horizon models shine.&lt;/p&gt;
&lt;p&gt;Examples from my own testing, with all the usual caveats: Opus-4.5 (Anthropic), GPT-5.2 (OpenAI), and GLM-4.7 (z.ai) have been solid for the long-horizon, get-it-moving-fast mode. Minimax M 2.1 (OpenCode Zen) feels closer to the short-horizon mode for me. Composer-1 (Cursor) also feels closer to that style. I suspect GPT-4o and GPT-5.1 (both OpenAI) might land there too, but I have not really test-driven them.&lt;/p&gt;
&lt;p&gt;The practical takeaway is that I now switch modes on purpose. When I need speed and initial momentum, I reach for long-horizon autonomy. When I need control, I choose a short-horizon model so I can babysit the work, watch the traces, and intercept it when it tries to do something clever in the wrong place.&lt;/p&gt;
&lt;h3 id="the-harness-lesson-cypress-beat-model-hopping"&gt;The harness lesson (Cypress beat model hopping)&lt;/h3&gt;&lt;p&gt;One more lesson from that period: I did a bunch of model hopping, trying to find something that would fix a particular class of behavioral breakage.&lt;/p&gt;
&lt;p&gt;The most frustrating failures were not subtle logic bugs. They were basic syntax errors introduced during tool-call patching, unclosed brackets, unclosed parentheses, that kind of thing. When that happens, you do not get a slightly-wrong feature, you get a page that fails to load. Debugging it manually is fine the first time, and infuriating on the seventh.&lt;/p&gt;
&lt;p&gt;The thing that actually moved the needle was listening to my colleague Anand Murthy and instantiating Cypress tests. A simple automated page reload catches those failures immediately. It shifts the pain earlier, gives the agent a verification loop it can run on demand, and turns "agentic coding" into something I can trust.&lt;/p&gt;
&lt;p&gt;Here is the dumbest possible example, taken straight from the Cypress suite in canvas-chat. It is not fancy, and that is the point. It catches "the page does not load" failures quickly.&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nx"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;Help Modal and Auto-Layout&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="nx"&gt;beforeEach&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;clearLocalStorage&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;clearIndexedDB&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;visit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="nx"&gt;it&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;opens and closes help modal&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;#help-btn&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;#help-modal&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;should&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;be.visible&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;#help-close&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;#help-modal&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;should&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;not.be.visible&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Beyond model choice, a great agentic harness matters. If your harness includes tests that a coding agent can run to verify no behavioral breakage, you get to move faster with more confidence, regardless of which model you are using.&lt;/p&gt;
&lt;h2 id="verbosity-attitude-and-the-cost-of-being-wrong"&gt;Verbosity, attitude, and the cost of being wrong&lt;/h2&gt;&lt;p&gt;The other axis that became obvious once I started pressure testing models is verbosity and its associated feel.&lt;/p&gt;
&lt;p&gt;I tried Gemini 2.5, and it was a disaster for me. After experiencing the long-horizon and short-horizon styles, I did not want to use it. It made elementary mistakes, like leaving trailing dangling curly braces where they were not supposed to be. Then it would apologize profusely over and over, like a Canadian on steroids. (I'm a born and bred Canadian, I'm allowed to say that!)&lt;/p&gt;
&lt;p&gt;In contrast, Claude and Opus are consistently upbeat and positive, and the same can be said for Minimax-M.2 and GLM-4.7. That matters more than I expected. When something breaks and you are iterating quickly, a model that stays constructive keeps the whole loop feeling fun.&lt;/p&gt;
&lt;p&gt;On the other end, GPT-5.2 would just go ahead and do things without being overly effusive, then loop back to tell me what it did. That sounds fine on paper, but it left me feeling a bit clueless. I would wonder what it was doing and whether I could intercept it if it went off on the wrong tangent. I often could not, because I needed to wait until the end to learn what it decided to do.&lt;/p&gt;
&lt;p&gt;So yes, I care about correctness. But I also care about how a model behaves while it is getting to correctness. The journey matters because the journey is where you spend your time.&lt;/p&gt;
&lt;h2 id="enthusiasm-is-a-feature"&gt;Enthusiasm is a feature&lt;/h2&gt;&lt;p&gt;This ties nicely to &lt;a href="https://x.com/Grady_Booch/status/2013343499563999589"&gt;a tweet&lt;/a&gt; I saw from Grady Booch (@Grady_Booch):&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;"The greatest value such tools have offered me is to reduce my cognitive load and automate various tedious tasks."&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here is the punchline:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;"To serve as an enthusiastic and indefatigable, albeit very naive and often unreliable, pair programmer."&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That enthusiasm and indefatigability, compared to a grumpy human, keeps the loop moving.&lt;/p&gt;
&lt;p&gt;My frustration pair coding with Gemini was not just the mistakes, it was the &lt;em&gt;emotional texture&lt;/em&gt; of the interaction. It would make mistakes and then apologize, repeatedly. After a while, you start optimizing your own behavior around the assistant's vibe, and that is not where you want your attention to go.&lt;/p&gt;
&lt;p&gt;A better pair programmer, human or AI, is relentlessly game for the next challenge. It affirms what you are trying to do, it corrects you when you are wrong, and it does not act like it is giving up. When the assistant stays constructive, the work stays fun.&lt;/p&gt;
&lt;h2 id="streaming-and-the-illusion-of-speed-the-harness-model"&gt;Streaming and the illusion of speed (the harness + model)&lt;/h2&gt;&lt;p&gt;Raw latency is one thing. What you &lt;em&gt;see&lt;/em&gt; while the model is working is another, and that is determined by the harness. In Cursor, you get a fast stream of tool calls and edits. You see the so-called thinking process. Something is clearly happening. With GLM-4.7 or Open Code in certain setups, you wait a long time with nothing streamed in—just a spinner or a blank state until the full response lands. Same model capability, same task, different harness, totally different experience. The harness that gives you a live trace makes the wait feel shorter and keeps you in the loop. The one that hides progress makes every request feel like a gamble. If you care about flow, streaming and visual feedback are not polish; they are table stakes, and they live in the harness.&lt;/p&gt;
&lt;h2 id="the-feel-is-also-vendor-lock-in"&gt;The "feel" is also vendor lock-in&lt;/h2&gt;&lt;p&gt;After enough hours with a single model, you start building muscle memory for its quirks. You learn how to phrase prompts so it does the right thing. You learn which mistakes to expect. You even learn its tone. That comfort is sticky.&lt;/p&gt;
&lt;p&gt;The sticky part is the problem. Getting used to a model's ergonomics is a form of vendor lock-in, and it is something I am determined to avoid.&lt;/p&gt;
&lt;p&gt;That is one reason I have been bouncing between models (apart from me hitting limits) to feel out the ragged frontier of model behavior. It is pretty revealing. You quickly learn that "best model" is not a single number. The model you want depends on whether you are scaffolding, refactoring, debugging, or doing the last-mile polish.&lt;/p&gt;
&lt;p&gt;If you want to keep your agency while using these tools, stay fluent across multiple feels. Otherwise you end up optimizing your workflow around one model's quirks and calling it productivity.&lt;/p&gt;
&lt;h2 id="a-more-pragmatic-way-to-think-about-model-choice"&gt;A more pragmatic way to think about model choice&lt;/h2&gt;&lt;p&gt;What I do now is less romantic than "find the best model". I think in terms of work phases and feedback loops.&lt;/p&gt;
&lt;p&gt;If I am scaffolding, I will happily take Opus-feel: longer-horizon autonomy and a big blob of output, because the cost of being wrong is usually low.&lt;/p&gt;
&lt;p&gt;If I am refactoring or debugging, I want Sonnet-feel: short-horizon iteration and tight supervision, because the cost of being wrong is a broken app and a bunch of time lost to verification.&lt;/p&gt;
&lt;p&gt;And if I keep hitting the same dumb failures, I try to fix my harness before I try to fix my model. Add the smallest test that fails fast, make it runnable by the agent, and suddenly the whole system behaves better. Cypress reloading the page and clicking one button did more for my sanity than another week of model hopping.&lt;/p&gt;
&lt;p&gt;At a systems level, I want a workflow where models are swappable components. In practice that means traces you can read, tests you can run, and a loop that tells you quickly when the agent broke something.&lt;/p&gt;
</content></entry><entry><title>How to build self-improving coding agents - Part 3</title><link href="https://ericmjl.github.io/blog/2026/1/19/how-to-build-self-improving-coding-agents-part-3/" rel="alternate"/><updated>2026-01-19T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:f1f0bd53-39f6-3f87-96d4-03384d798a21</id><content type="html">&lt;p&gt;In &lt;a href="../../17/how-to-build-self-improving-coding-agents-part-1/"&gt;part 1&lt;/a&gt;, I covered &lt;code&gt;AGENTS.md&lt;/code&gt; as repo memory.&lt;/p&gt;
&lt;p&gt;In &lt;a href="../../18/how-to-build-self-improving-coding-agents-part-2/"&gt;part 2&lt;/a&gt;, I covered skills as reusable playbooks.&lt;/p&gt;
&lt;p&gt;This post is about turning those two ideas into something you can run as a practice.&lt;/p&gt;
&lt;h2 id="the-maturity-model"&gt;The maturity model&lt;/h2&gt;&lt;p&gt;Once you have both repo memory and skills, you can think about how the practice evolves over time.&lt;/p&gt;
&lt;h3 id="stage-0-ad-hoc-prompting"&gt;Stage 0: Ad hoc prompting&lt;/h3&gt;&lt;p&gt;You keep re-explaining the same things in chat. It works, but it does not compound.&lt;/p&gt;
&lt;h3 id="stage-1-repo-local-memory"&gt;Stage 1: Repo-local memory&lt;/h3&gt;&lt;p&gt;You add repository-specific guardrails and a code map.&lt;/p&gt;
&lt;p&gt;This is where &lt;code&gt;AGENTS.md&lt;/code&gt; shines.&lt;/p&gt;
&lt;h3 id="stage-2-global-personal-skills"&gt;Stage 2: Global personal skills&lt;/h3&gt;&lt;p&gt;Once a workflow repeats across repos, you promote it into a global skill on your machine.&lt;/p&gt;
&lt;p&gt;If you want a concrete bootstrap set, here is what I would install globally:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/ericmjl/skills/blob/main/skills/skill-creator/"&gt;&lt;code&gt;skill-creator&lt;/code&gt;&lt;/a&gt;: lowers the activation energy for making new skills.&lt;/li&gt;
&lt;li&gt;an installer and updater for skills, for example &lt;a href="https://github.com/numman-ali/openskills"&gt;&lt;code&gt;openskills&lt;/code&gt;&lt;/a&gt;: makes distribution and updates less annoying.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ericmjl/skills/tree/main/skills/agents-md-improver"&gt;&lt;code&gt;agents-md-improver&lt;/code&gt;&lt;/a&gt;: keeps the repo map current without you thinking about it.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="stage-3-shared-skills"&gt;Stage 3: Shared skills&lt;/h3&gt;&lt;p&gt;If a workflow repeats across a team, it belongs in a shared location with a clear install path.&lt;/p&gt;
&lt;p&gt;I do not think you should start here. Start repo-local, then promote only when you feel the pain twice.&lt;/p&gt;
&lt;p&gt;Promotion decisions come from paying attention to what the agent actually does in practice.&lt;/p&gt;
&lt;h2 id="watch-traces-then-distill-constraints"&gt;Watch traces, then distill constraints&lt;/h2&gt;&lt;p&gt;If you work with agents long enough, you start to notice the model’s default moves.&lt;/p&gt;
&lt;p&gt;When I see an agent repeatedly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;taking an overcomplicated path&lt;/li&gt;
&lt;li&gt;missing a file I know is relevant&lt;/li&gt;
&lt;li&gt;applying a global refactor when a surgical fix is needed&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I treat that as a signal.&lt;/p&gt;
&lt;p&gt;Then I decide what kind of fix it is.&lt;/p&gt;
&lt;p&gt;If it is a repo invariant, a navigation hint, or a local norm, it belongs in &lt;code&gt;AGENTS.md&lt;/code&gt;. That is the always-on context for how work should happen in this repo.&lt;/p&gt;
&lt;p&gt;If it is a repeatable procedure with a clear output contract, it belongs in a skill.&lt;/p&gt;
&lt;p&gt;Sometimes the procedure is repo-specific. In that case I keep it as a repo-local skill. If I feel the pain twice in another repo, I promote it into a global skill.&lt;/p&gt;
&lt;p&gt;This is how you get operational learning without pretending the model is learning.&lt;/p&gt;
&lt;p&gt;Underneath, a lot of this comes down to writing instructions in a way that can be executed.&lt;/p&gt;
&lt;h2 id="markdown-is-becoming-executable"&gt;Markdown is becoming executable&lt;/h2&gt;&lt;p&gt;One reason this whole approach works is that the agent can execute what you write.&lt;/p&gt;
&lt;p&gt;When an LLM can execute tool calls, Markdown becomes an executable language.&lt;/p&gt;
&lt;p&gt;Skills fit this pattern. A &lt;code&gt;SKILL.md&lt;/code&gt; is just a structured instruction sheet, but it is also runnable in the sense that the agent can turn it into searches, file reads, edits, and command execution.&lt;/p&gt;
&lt;p&gt;The other trick is that skills are loaded on demand. The agent reads a short description first, then loads the full instructions only when it needs them.&lt;/p&gt;
&lt;p&gt;You can write a precise plan in plain language, and the agent can turn it into:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;searches&lt;/li&gt;
&lt;li&gt;file reads&lt;/li&gt;
&lt;li&gt;surgical edits&lt;/li&gt;
&lt;li&gt;test runs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is not magic. It still depends on linguistic precision. But the ergonomics shift. You can describe a workflow at the level you actually think about it, then let the agent do the clerical work.&lt;/p&gt;
&lt;p&gt;This is also why I like the runbook analogy, even with the caveats.&lt;/p&gt;
&lt;h2 id="when-to-update-agents-md-vs-create-a-skill"&gt;When to update &lt;code&gt;AGENTS.md&lt;/code&gt; vs create a skill&lt;/h2&gt;&lt;p&gt;Skills tell an agent how to do something.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; tells an agent how this repo works, and what rules it must follow while doing anything at all.&lt;/p&gt;
&lt;p&gt;Here is how I decide.&lt;/p&gt;
&lt;p&gt;Update &lt;code&gt;AGENTS.md&lt;/code&gt; when the instruction is specific to the repo:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;navigation help: where things live, what files matter, what to ignore&lt;/li&gt;
&lt;li&gt;local norms: build commands, test commands, environment rules, style constraints&lt;/li&gt;
&lt;li&gt;guardrails: what not to do in this repo&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Create a skill when the workflow is reusable, or when you want a named, on-demand playbook:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a multi-step procedure you want to invoke repeatedly&lt;/li&gt;
&lt;li&gt;a workflow that spans repos or products&lt;/li&gt;
&lt;li&gt;a task with a strict output contract (release announcements, status updates, summaries)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If I am unsure, I start repo-local. If I feel the pain twice in another repo, I promote it into a global skill.&lt;/p&gt;
&lt;h2 id="the-meta-skill-is-metacognition"&gt;The meta skill is metacognition&lt;/h2&gt;&lt;p&gt;The most valuable “skill”, however, is not a file format. It is the habit of watching yourself work.&lt;/p&gt;
&lt;p&gt;I try to ask: what am I doing repeatedly that should be systematized?&lt;/p&gt;
&lt;p&gt;If the answer is “I keep re-explaining how this repo is organized”, that goes into &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;If the answer is “I keep asking for the same kind of summary, debug sequence, or release note format”, that becomes a skill.&lt;/p&gt;
&lt;p&gt;Once you start doing this, you build a compounding loop. The agent handles more of the repeated work, and you spend more time on judgment and design.&lt;/p&gt;
&lt;p&gt;If this all sounds like more than coding, that is because it is.&lt;/p&gt;
&lt;h2 id="where-this-seems-to-be-going"&gt;Where this seems to be going&lt;/h2&gt;&lt;p&gt;I buy Simon Willison’s framing that these tools are general agents disguised as developer tools (&lt;a href="https://simonwillison.net/2026/Jan/12/claude-cowork/"&gt;Claude Cowork post&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Even if you start with coding, the moment an agent can run terminal commands and manipulate files, the surface area expands to “almost anything”, as long as you know how to steer it.&lt;/p&gt;
&lt;p&gt;That matches how I use coding agents.&lt;/p&gt;
&lt;p&gt;Yes, I use it for coding work. But I also use it for other intellectual work: ghostwriting blog posts (which I scrutinize heavily, because the review process is essential for me to own the content), writing release announcements, and turning messy notes into structured drafts.&lt;/p&gt;
&lt;p&gt;I have also heard Theo Brown make a similar point when talking about Claude Cowork (&lt;a href="https://www.youtube.com/watch?v=IcQEaopx90g"&gt;video&lt;/a&gt;). The details vary, but the pattern is the same: once you have a general agent, the label “coding tool” becomes more about marketing and UI than capability.&lt;/p&gt;
&lt;p&gt;So I am increasingly convinced that the long-term shape here is web-deployed agents with less scary branding.&lt;/p&gt;
&lt;p&gt;You will still want composable components for LLM workflows. But for day-to-day work, the most useful thing is an agent that can execute commands and apply changes, while carrying a growing set of skills and repository memory.&lt;/p&gt;
&lt;p&gt;That combination is what makes the agent feel less like a chat box and more like a teammate.&lt;/p&gt;
</content></entry><entry><title>How to build self-improving coding agents - Part 2</title><link href="https://ericmjl.github.io/blog/2026/1/18/how-to-build-self-improving-coding-agents-part-2/" rel="alternate"/><updated>2026-01-18T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:8caf7de9-77ba-3b15-be93-8de346553886</id><content type="html">&lt;p&gt;In &lt;a href="../../17/how-to-build-self-improving-coding-agents-part-1/"&gt;part 1&lt;/a&gt;, I focused on repo memory with &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;In this post, I am switching to the other lever: skills.&lt;/p&gt;
&lt;h2 id="skills-are-prompt-compression"&gt;Skills are prompt compression&lt;/h2&gt;&lt;p&gt;Skills are the other half of the system.&lt;/p&gt;
&lt;p&gt;When a task repeats, I do not want to keep re-explaining the workflow. I want a playbook I can invoke.&lt;/p&gt;
&lt;h3 id="what-a-skill-is"&gt;What a skill is&lt;/h3&gt;&lt;p&gt;A skill is a folder with a &lt;code&gt;SKILL.md&lt;/code&gt; file.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;SKILL.md&lt;/code&gt; is the prompt. The bundled scripts and assets are the tool layer.&lt;/p&gt;
&lt;p&gt;A good skill makes three things explicit:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;when to use it&lt;/li&gt;
&lt;li&gt;what steps to take&lt;/li&gt;
&lt;li&gt;what good output looks like&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you want the spec, see &lt;a href="https://agentskills.io/home"&gt;Agent Skills&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Skills are best formed around jobs to be done: concrete, repeatable workflows rather than abstract capabilities. Think "debug a GitHub Actions failure" or "draft a release announcement," not "know about CI" or "write good prose." When the job is clear, the skill has a natural boundary and a clear trigger. When it is vague, the skill is hard to invoke and hard to improve.&lt;/p&gt;
&lt;p&gt;A wrong framing is "skills for tools." Skills get invoked in the loop of trying to accomplish a job, not in the context of trying to use a tool. The tool is a means; the job is why you reach for it. If you design a skill around a tool, you end up with something the agent has to remember to use. If you design it around a job, the agent reaches for it when the job shows up.&lt;/p&gt;
&lt;h3 id="examples"&gt;Examples&lt;/h3&gt;&lt;p&gt;A GitHub debugging skill is the obvious starting point. CI failures are repetitive and usually want the same sequence: identify failing jobs, pull logs, inspect diffs, reproduce locally, then patch.&lt;/p&gt;
&lt;p&gt;A second example is a release announcement skill.&lt;/p&gt;
&lt;p&gt;The motivation here was not abstract. I was spending a good half hour each release just trying to compose the announcement, and I did not want to do that anymore.&lt;/p&gt;
&lt;p&gt;The output contract was also specific. I wanted release announcements that are copy-pasteable into Microsoft Teams, with emojis, but otherwise minimal formatting because Teams formatting is inconsistent.&lt;/p&gt;
&lt;p&gt;A third example is more technical.&lt;/p&gt;
&lt;p&gt;At work I had a session with a coding agent to train an ML model inside a script. After that session, I had it write a report on what it learned and what changed. Then I turned that report writing into a skill.&lt;/p&gt;
&lt;p&gt;The report format was familiar to everyone on the team: Abstract, Introduction, Methods, Results, Discussion.&lt;/p&gt;
&lt;p&gt;The content came from real artifacts: stdout logs, metrics, code, config files, git diffs, and the agent’s own session history.&lt;/p&gt;
&lt;p&gt;A fourth example is about tacit domain expertise.&lt;/p&gt;
&lt;p&gt;A teammate of mine created a skill that encoded her implicit knowledge from years of debugging chromatography traces. The point was not that the agent suddenly became a scientist. The point was that her debugging procedure became explicit and reusable.&lt;/p&gt;
&lt;h3 id="skill-creation-and-iteration"&gt;Skill creation and iteration&lt;/h3&gt;&lt;p&gt;I now like skills because they are easy to iterate on. I used to be more skeptical, and I still think MCP servers have a cleaner distribution story, but my opinion has shifted as I have used skills more in real workflows (&lt;a href="https://ericmjl.github.io/blog/2025/10/20/exploring-skills-vs-mcp-servers/"&gt;Exploring Skills vs MCP Servers&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;For the release announcements, I fed my coding agent a few examples of what “good” looked like. I was using Anthropic’s &lt;a href="https://github.com/ericmjl/skills/blob/main/skills/skill-creator/"&gt;&lt;code&gt;skill-creator&lt;/code&gt;&lt;/a&gt; skill at the time, and those examples became part of the skill itself, stored as assets that the agent could reuse.&lt;/p&gt;
&lt;p&gt;This is a huge energy barrier reducer. It is much easier to iterate on a Markdown-based skill than it is to start from scratch with “write me a Python script that does X”. You can still add scripts inside a skill when you need determinism, but the interface is the Markdown.&lt;/p&gt;
&lt;p&gt;The other half is the feedback loop. When I edit the generated release announcement, I feed the revised version back to the agent and tell it to update the skill with the new example. That way the skill evolves as my taste evolves.&lt;/p&gt;
&lt;p&gt;This is also a way to share. A skill is reviewable. I can open a PR and let collaborators comment on both the output and the process that produced it.&lt;/p&gt;
&lt;p&gt;In the chromatography example, using &lt;a href="https://github.com/ericmjl/skills/blob/main/skills/skill-creator/"&gt;&lt;code&gt;skill-creator&lt;/code&gt;&lt;/a&gt; to generate the first draft mattered for another reason too. English is not my teammate’s first language. The structure makes it much easier to get from “I know what I do” to “here is the procedure an agent can follow”.&lt;/p&gt;
&lt;h3 id="distribution-and-updates"&gt;Distribution and updates&lt;/h3&gt;&lt;p&gt;This is where skills feel less mature than MCP servers.&lt;/p&gt;
&lt;p&gt;An MCP server has a clean distribution story. You can &lt;code&gt;pip install&lt;/code&gt; it, configure auth once, and you get a centrally versioned bundle of prompts and tools. Updating is a normal package update.&lt;/p&gt;
&lt;p&gt;Skills still involve moving folders between machines and repos, and remembering where each harness expects skills to live.&lt;/p&gt;
&lt;p&gt;I originally ended up writing a &lt;a href="https://github.com/ericmjl/skills/tree/main/skills/skill-installer"&gt;&lt;code&gt;skill-installer&lt;/code&gt;&lt;/a&gt; skill. It is the same move as &lt;a href="https://github.com/ericmjl/skills/blob/main/skills/skill-creator/"&gt;&lt;code&gt;skill-creator&lt;/code&gt;&lt;/a&gt;, but for distribution and updates.&lt;/p&gt;
&lt;p&gt;When I say “install this skill” or “update this skill from this URL”, the agent needs to ask two key questions if I have not already specified them:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;is this repo-local or machine-global?&lt;/li&gt;
&lt;li&gt;which harnesses should discover it?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then it does the boring part consistently.&lt;/p&gt;
&lt;p&gt;Update: it looks like &lt;a href="https://github.com/numman-ali/openskills"&gt;&lt;code&gt;openskills&lt;/code&gt;&lt;/a&gt; now solves most of what I wanted here, and it does it more deterministically. It is a CLI that installs skill folders from GitHub or local paths, tracks their sources for updates, and can target multiple install locations.&lt;/p&gt;
&lt;p&gt;OpenSkills has a "universal" mode that installs to &lt;code&gt;.agent/skills&lt;/code&gt; (repo) and &lt;code&gt;~/.agent/skills&lt;/code&gt; (machine).&lt;/p&gt;
&lt;p&gt;The caveat is that &lt;code&gt;.agent/skills&lt;/code&gt; is not a universal discovery standard across harnesses. Some tools look in &lt;code&gt;.claude/skills&lt;/code&gt;, &lt;code&gt;.github/skills&lt;/code&gt;, &lt;code&gt;.opencode&lt;/code&gt;, or other locations. So OpenSkills helps with deterministic installs and updates, but you still need to know what your harness will actually read.&lt;/p&gt;
&lt;p&gt;I expect this to converge soon.&lt;/p&gt;
&lt;p&gt;At this point you have both memory and playbooks. The question becomes how you decide what to invest in next.&lt;/p&gt;
&lt;h2 id="coming-next"&gt;Coming next&lt;/h2&gt;&lt;p&gt;Part 3 covers the operating model.&lt;/p&gt;
&lt;p&gt;It lays out a maturity model, a concrete bootstrap set of skills to install globally, and a decision rule for when to update &lt;code&gt;AGENTS.md&lt;/code&gt; versus when to create a skill.&lt;/p&gt;
&lt;p&gt;&lt;a href="../../19/how-to-build-self-improving-coding-agents-part-3/"&gt;How to build self-improving coding agents - Part 3&lt;/a&gt;&lt;/p&gt;
</content></entry><entry><title>How to build self-improving coding agents - Part 1</title><link href="https://ericmjl.github.io/blog/2026/1/17/how-to-build-self-improving-coding-agents-part-1/" rel="alternate"/><updated>2026-01-17T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:f5c71a26-6381-3980-aa4e-20d1deeb4eb3</id><content type="html">&lt;p&gt;I want my coding agents to get better every week.&lt;/p&gt;
&lt;p&gt;Not in the abstract “the models are improving” sense. I mean it in the operational sense: if an agent makes a mistake, or takes a path I would not take, I want that feedback to stick. If I have to repeat the same preference every session, I am not using an agent. I am babysitting a very fast intern.&lt;/p&gt;
&lt;p&gt;The trick is that the model weights are not changing mid-week. So if you want “self-improvement”, you need to change the environment the agent works inside.&lt;/p&gt;
&lt;p&gt;I have found two levers that compound:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; as repository memory&lt;/li&gt;
&lt;li&gt;skills as reusable playbooks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This post is a longer “source of truth” version. My intent is to later break it into smaller blog entries, and also rework it into chapters for my data science bootstrap notes.&lt;/p&gt;
&lt;h2 id="where-improvement-comes-from"&gt;Where improvement comes from&lt;/h2&gt;&lt;p&gt;The UX I am after is simple: I stop repeating myself. I stop doing the same end-of-day cleanup, writing the same reminders, re-explaining where files live. The agent starts each session closer to how I want it to work.&lt;/p&gt;
&lt;p&gt;If the model weights are not changing mid-week, improvement has to come from the environment you wrap around the agent.&lt;/p&gt;
&lt;p&gt;For me that environment has two pieces:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;durable repository memory (&lt;code&gt;AGENTS.md&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;reusable playbooks (skills)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Once you have those two, you can treat “agent improvement” like runbooks plus postmortems.&lt;/p&gt;
&lt;p&gt;The analogy is imperfect, because this is not documentation for humans. The loop is the same though: write down the repeatable steps, then write down what surprised you and what you will do differently next time.&lt;/p&gt;
&lt;p&gt;The difference is that natural language can turn into tool calls. When you write things down precisely, the agent can execute them.&lt;/p&gt;
&lt;p&gt;I usually start with &lt;code&gt;AGENTS.md&lt;/code&gt;, because it cuts down exploration immediately.&lt;/p&gt;
&lt;h2 id="agents-md-as-repository-memory"&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; as repository memory&lt;/h2&gt;&lt;p&gt;If you have not run into the &lt;code&gt;AGENTS.md&lt;/code&gt; convention before, see &lt;a href="https://agents.md/"&gt;agents.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;To be effective, &lt;code&gt;AGENTS.md&lt;/code&gt; needs to do two things for the agent.&lt;/p&gt;
&lt;p&gt;First, it needs to make the agent fast at navigating the repo so it can get to the right files with minimal wandering. A code map is a straightforward way to do that.&lt;/p&gt;
&lt;p&gt;Second, it needs to encode the local ways of working in this repo so the agent stops repeating the same mistakes. That is where corrections and norms live.&lt;/p&gt;
&lt;p&gt;This is the loop I want:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I observe a mismatch.&lt;/li&gt;
&lt;li&gt;I tell the agent what must be true.&lt;/li&gt;
&lt;li&gt;The agent writes the correction into &lt;code&gt;AGENTS.md&lt;/code&gt; (or a repo-local skill).&lt;/li&gt;
&lt;li&gt;The agent reads it next time.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="fast-navigation-to-the-right-files"&gt;Fast navigation to the right files&lt;/h3&gt;&lt;p&gt;In the ideal state, the agent gets to the right files quickly.&lt;/p&gt;
&lt;p&gt;A code map is the simplest way I know to make that happen. It does not have to be perfect. It can be slightly stale and still be useful.&lt;/p&gt;
&lt;p&gt;I have seen this pay off in a very practical way. In my &lt;code&gt;canvas-chat&lt;/code&gt; codebase, having a map of the repo let the agent one-shot an obscure spot where events were emitted for node rendering. Without a map, the agent previously needed 5 to 6 &lt;code&gt;rg&lt;/code&gt; searches, just to find the right neighborhood of the code.&lt;/p&gt;
&lt;p&gt;The difference is small in absolute time, something like 40 seconds versus 2 seconds. But it changes the feel of the collaboration. The agent spends less time wandering, and I spend less time steering.&lt;/p&gt;
&lt;h3 id="close-the-loop-when-the-map-is-stale"&gt;Close the loop when the map is stale&lt;/h3&gt;&lt;p&gt;There is one extra move that makes this feel self-correcting: When the agent notices that the code map looks stale, it should update the code map.&lt;/p&gt;
&lt;p&gt;This is a subtle point. The map is not a static artifact. It is part of a feedback loop. When the agent’s exploration discovers a mismatch between the map and reality, that discovery should flow back into &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;You can encode this as an explicit instruction inside &lt;code&gt;AGENTS.md&lt;/code&gt;. You can also refresh on a schedule, like weekly, but the on-demand update is the part that makes the loop feel alive.&lt;/p&gt;
&lt;h3 id="corrections-that-become-durable-norms"&gt;Corrections that become durable norms&lt;/h3&gt;&lt;p&gt;The second job of &lt;code&gt;AGENTS.md&lt;/code&gt; is to hold repo-specific corrections to agents behaviour.&lt;/p&gt;
&lt;p&gt;These are the things you find yourself saying out loud.&lt;/p&gt;
&lt;p&gt;Two examples from my own work:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Run Python in the &lt;code&gt;pixi&lt;/code&gt; context. Use &lt;code&gt;pixi run python ...&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Do not cheat by modifying the tests to make them pass.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I say the first one because the agent will often try &lt;code&gt;python -c ...&lt;/code&gt; to quickly check something. In a &lt;code&gt;pixi&lt;/code&gt;-managed project, that fails if you do not have a global Python.&lt;/p&gt;
&lt;p&gt;I say the second one because changing tests to make them pass destroys the point of having tests.&lt;/p&gt;
&lt;p&gt;Once these rules are written down, the agent stops making you restate them. This is the simplest way I know to reduce repeated friction.&lt;/p&gt;
&lt;h2 id="a-starter-prompt-for-generating-agents.md`"&gt;A starter prompt for generating &lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/h2&gt;&lt;p&gt;I have found it useful to bootstrap &lt;code&gt;AGENTS.md&lt;/code&gt; with a one-time deep dive.&lt;/p&gt;
&lt;p&gt;Here is a prompt I use as a starting point. It is intentionally repo-specific.&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;You are a coding agent. Read through this repository and create an `AGENTS.md` file at the repo root.

Requirements:
- Include a short codebase map that helps an agent find files quickly.
- Focus on entry points, directory roles, naming conventions, configuration wiring, and test locations.
- Add a section called &amp;quot;Local norms&amp;quot; with repo-specific rules you infer from the code and tooling.
- Add a section called &amp;quot;Self-correction&amp;quot; with two explicit instructions:
  - If the code map is discovered to be stale, update it.
  - If the user gives a correction about how work should be done in this repo, add it to &amp;quot;Local norms&amp;quot; (or another clearly labeled section) so future sessions inherit it.

Process:
- Use search and targeted file reads, do not read every file.
- Prefer `rg` searches to find entry points and configs.
- Prefer high-signal files: `README`, `pyproject.toml`, `package.json`, `Makefile`, `opencode.json`, `.github/workflows`, and top-level `src` or `app` directories.

Output:
- Write the final `AGENTS.md` contents in Markdown.
- Keep it concise. Optimize for navigation and correctness.
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If you want, you can go further and add a cadence rule like “refresh weekly”, but I would keep it lightweight. The goal is compounding value, not bureaucracy.&lt;/p&gt;
&lt;p&gt;Once &lt;code&gt;AGENTS.md&lt;/code&gt; exists, skills are the second lever.&lt;/p&gt;
&lt;h2 id="coming-next"&gt;Coming next&lt;/h2&gt;&lt;p&gt;Part 2 is about skills as reusable playbooks.&lt;/p&gt;
&lt;p&gt;It covers what a skill is, several examples from coding and scientific work, and why I ended up writing a &lt;code&gt;skill-installer&lt;/code&gt; skill to deal with the current distribution story.&lt;/p&gt;
&lt;p&gt;&lt;a href="../../18/how-to-build-self-improving-coding-agents-part-2/"&gt;How to build self-improving coding agents - Part 2&lt;/a&gt;&lt;/p&gt;
</content></entry><entry><title>How I fixed a browser selection bug with sequence alignment algorithms</title><link href="https://ericmjl.github.io/blog/2026/1/6/how-i-fixed-a-browser-selection-bug-with-sequence-alignment-algorithms/" rel="alternate"/><updated>2026-01-06T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:8ff1d33b-f8e0-3e00-bb6e-75771bc021b1</id><content type="html">&lt;p&gt;I ran into a frustrating bug this week in &lt;a href="https://github.com/ericmjl/canvas-chat"&gt;canvas-chat&lt;/a&gt;, my experimental canvas-based chat interface &lt;a href="../../../../2025/12/31/canvas-chat-a-visual-interface-for-thinking-with-llms/"&gt;I built at the end of last year&lt;/a&gt;. The bug seemed simple on the surface: when users selected text from a rendered markdown table and clicked to highlight it, the highlighting would sometimes stop partway through, or highlight the wrong characters entirely.&lt;/p&gt;
&lt;p&gt;What started as a "quick fix" turned into a journey through several failed approaches before I remembered an algorithm from my undergraduate bioinformatics days. Sometimes the best solution to a problem comes from an unexpected domain.&lt;/p&gt;
&lt;h2 id="the-problem-browser-selections-are-messy"&gt;The problem: Browser selections are messy&lt;/h2&gt;&lt;p&gt;Canvas-chat has a feature where you can select text from an AI response, and the app creates a "highlight" node that links back to the source. When you click on the highlight, the corresponding text in the source gets wrapped in a &lt;code&gt;&amp;lt;mark&amp;gt;&lt;/code&gt; tag.&lt;/p&gt;
&lt;p&gt;This worked fine for simple paragraphs. But when I tried it on tables containing KaTeX-rendered math, things went wrong:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What I expected to highlight:&lt;/strong&gt; &lt;mark&gt;66.00 (0.18±0.58)&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What actually got highlighted:&lt;/strong&gt; &lt;mark&gt;66.00 ( 0.18&lt;/mark&gt;±0.58)&lt;/p&gt;
&lt;p&gt;The highlighting was off by more than a few characters, and would stop before the end of my selection. In some cases, it would highlight completely wrong sections.&lt;/p&gt;
&lt;h2 id="digging-into-the-root-cause"&gt;Digging into the root cause&lt;/h2&gt;&lt;p&gt;The problem came from how KaTeX renders math and how browsers handle text selection.&lt;/p&gt;
&lt;p&gt;KaTeX renders math with multiple text representations:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;katex&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="cm"&gt;&amp;lt;!-- MathML for accessibility/screen readers --&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;katex-mathml&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;math&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;mn&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;0.13&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;mn&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;±&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;mo&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;math&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="cm"&gt;&amp;lt;!-- Visual HTML for display --&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;katex-html&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;mord&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;0.13&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt; &lt;span class="na"&gt;class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;mord&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;±&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;span&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;When you select text that spans across KaTeX-rendered content, &lt;code&gt;selection.toString()&lt;/code&gt; gives you something like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;"66.00 (
0.13
±
0.13±0.58)"
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Notice the duplicated &lt;code&gt;0.13&lt;/code&gt; and the random newlines? The browser included text from both the MathML (for accessibility) and the visual spans. Add in tabs between table cells and inconsistent spacing around operators, and you have a string that looks nothing like the clean HTML text content.&lt;/p&gt;
&lt;h2 id="first-attempt-normalization-layers"&gt;First attempt: Normalization layers&lt;/h2&gt;&lt;p&gt;My initial approach was to normalize both strings (the user's selection and the HTML text) before matching:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Collapse all whitespace to single spaces&lt;/li&gt;
&lt;li&gt;Remove KaTeX duplication patterns (like &lt;code&gt;0.13 ± 0.13±&lt;/code&gt; → &lt;code&gt;0.13±&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Normalize spacing around &lt;code&gt;±&lt;/code&gt; operators&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Then find the match in the normalized strings, and map the positions back to the original.&lt;/p&gt;
&lt;p&gt;This is where things got complicated. I needed to track:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Which positions in the normalized string corresponded to which positions in the original&lt;/li&gt;
&lt;li&gt;How to reverse the mapping after finding a match&lt;/li&gt;
&lt;li&gt;How to handle characters that got removed entirely during normalization&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The code became a tangled mess of position arrays and off-by-one bugs. Here's a simplified version of what it looked like:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;// Build mapping from normalized to original positions&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;normalizedToOriginal&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;inWhitespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;leadingTrimmed&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;fullText&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;ch&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;fullText&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/\s/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ch&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;inWhitespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;leadingTrimmed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="nx"&gt;normalizedToOriginal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;inWhitespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;leadingTrimmed&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;inWhitespace&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nx"&gt;normalizedToOriginal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Then also account for math spacing normalization...&lt;/span&gt;
&lt;span class="c1"&gt;// And KaTeX deduplication...&lt;/span&gt;
&lt;span class="c1"&gt;// Each layer compounds the position mapping complexity&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The position mapping kept breaking. I'd fix one case only to break another. &lt;strong&gt;I was trying to maintain a bijection between two strings that had been transformed through multiple non-invertible operations.&lt;/strong&gt; It wasn't going to work.&lt;/p&gt;
&lt;h2 id="the-insight-this-is-a-sequence-alignment-problem"&gt;The insight: This is a sequence alignment problem&lt;/h2&gt;&lt;p&gt;After banging my head against the normalization approach for a while, I took a step back. What was I actually trying to do?&lt;/p&gt;
&lt;p&gt;I had two strings:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The user's selection (messy, with artifacts)&lt;/li&gt;
&lt;li&gt;The HTML text content (clean)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I needed to find where the user's selection "matched" in the HTML text, tolerating:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Insertions (extra whitespace, duplicated characters in the selection)&lt;/li&gt;
&lt;li&gt;Deletions (characters present in HTML but not in selection)&lt;/li&gt;
&lt;li&gt;Mismatches (different whitespace characters)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is exactly what sequence alignment algorithms are designed for. In bioinformatics, we use these algorithms to compare DNA or protein sequences that may have evolved with insertions, deletions, and mutations. The classic algorithm for finding the best local alignment between two sequences is &lt;strong&gt;Smith-Waterman&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;I learned Smith-Waterman as an undergraduate, probably around 2008. I never thought I'd use it for web development.&lt;/p&gt;
&lt;h2 id="the-solution-align-the-beginning-and-end"&gt;The solution: Align the beginning and end&lt;/h2&gt;&lt;p&gt;I didn't need to align the entire selection - I just needed to find where it started and ended in the HTML text. So I:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Take the first ~20 characters of the user's selection and align them to find the &lt;strong&gt;start position&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Take the last ~20 characters, reverse both strings, align to find the &lt;strong&gt;end position&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Here's the core alignment function:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kd"&gt;function&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;alignStart&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;queryPrefix&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;queryPrefix&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;MATCH&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="c1"&gt;// Reward for matching characters&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;MISMATCH&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="c1"&gt;// Penalty for different characters&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;GAP&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt;       &lt;/span&gt;&lt;span class="c1"&gt;// Penalty for insertions/deletions&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;WS_MATCH&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="c1"&gt;// Softer reward for whitespace matches&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="c1"&gt;// Build the scoring matrix&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;=&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;maxScore&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;maxI&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;maxJ&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;qChar&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;queryPrefix&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;tChar&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;matchVal&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;qChar&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;===&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;tChar&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;matchVal&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="sr"&gt;/\s/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;qChar&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;WS_MATCH&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;MATCH&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/\s/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;qChar&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="sr"&gt;/\s/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tChar&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;matchVal&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;WS_MATCH&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="c1"&gt;// Any whitespace matches any whitespace&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;matchVal&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;MISMATCH&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="mf"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="c1"&gt;// Local alignment can restart anywhere&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;matchVal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="c1"&gt;// Diagonal: match/mismatch&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;GAP&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;          &lt;/span&gt;&lt;span class="c1"&gt;// Up: gap in target&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;GAP&lt;/span&gt;&lt;span class="w"&gt;           &lt;/span&gt;&lt;span class="c1"&gt;// Left: gap in query&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;maxScore&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;maxScore&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;maxI&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;                &lt;/span&gt;&lt;span class="nx"&gt;maxJ&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="c1"&gt;// Traceback to find start position&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="c1"&gt;// ... (walk backwards from maxI, maxJ to find where alignment began)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The key insight is that Smith-Waterman's local alignment naturally handles all the messiness. Extra newlines in the selection? They're just gaps. Duplicated numbers? They align to the same position. Different whitespace characters? They all match each other.&lt;/p&gt;
&lt;h2 id="the-result"&gt;The result&lt;/h2&gt;&lt;p&gt;The new approach passes all the test cases that the normalization approach failed:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Test: Simple word&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Target: &lt;code&gt;"Hello world, this is a test."&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Query: &lt;code&gt;"world"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Result: &lt;mark&gt;world&lt;/mark&gt; (positions 6-11)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Test: KaTeX duplication&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Target: &lt;code&gt;"66.00 (0.18 ± 0.18±0.58)"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Query: &lt;code&gt;"66.00 (\n0.18\n±\n0.18±0.58)"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Result: &lt;mark&gt;66.00 (0.18 ± 0.18±0.58)&lt;/mark&gt; (positions 0-25)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Test: Cross-block selection&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Target: &lt;code&gt;"The Heading Some paragraph text here."&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Query: &lt;code&gt;"The Heading\n\nSome paragraph"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Result: &lt;mark&gt;The Heading Some paragraph&lt;/mark&gt; (positions 0-25)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-lesson-know-your-algorithms"&gt;The lesson: Know your algorithms&lt;/h2&gt;&lt;p&gt;I didn't invent anything new here. Smith-Waterman has been around since 1981. I just recognized that my web development problem was, at its core, a sequence alignment problem.&lt;/p&gt;
&lt;p&gt;This is why I think it's valuable to study algorithms and techniques from different domains, even if they seem unrelated to your day-to-day work. You never know when dynamic programming from bioinformatics will solve your JavaScript text highlighting bug.&lt;/p&gt;
&lt;p&gt;The normalization approach was trying to make two messy things identical before comparing them. The alignment approach embraced the messiness and asked: "Given that these are different, where do they best correspond?"&lt;/p&gt;
&lt;p&gt;That's a fundamentally different framing, and it's the framing that actually matched the problem.&lt;/p&gt;
&lt;p&gt;Interestingly, I couldn't find prior examples of using Smith-Waterman specifically for UI text highlighting or matching browser text selections to source HTML. The algorithm is well-established in bioinformatics for DNA and protein sequence alignment, and it appears in some fuzzy string matching contexts like spell-checking and record linkage. But applying it to handle the specific artifacts that browsers introduce when selecting text from rendered HTML with KaTeX, MathML, or complex table structures? That seems to be a new application. Sometimes the best solutions come from recognizing that your problem, despite appearing domain-specific, maps onto a well-solved problem from an entirely different field.&lt;/p&gt;
&lt;p&gt;One more note: I didn't write the JavaScript implementation myself. I directed Claude Opus 4.5 in &lt;a href="https://opencode.ai"&gt;OpenCode&lt;/a&gt; to write it for me. My contribution was recognizing that this was a sequence alignment problem and describing the approach - the actual code was generated by the AI. This is becoming my preferred way to work: I provide the domain insight and algorithmic direction, and the AI handles the implementation details.&lt;/p&gt;
&lt;h2 id="appendix-the-full-solution"&gt;Appendix: The full solution&lt;/h2&gt;&lt;p&gt;For those curious, the complete implementation is in the &lt;a href="https://github.com/ericmjl/canvas-chat/pull/97"&gt;pull request&lt;/a&gt;. The key functions are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;alignStart(queryPrefix, target)&lt;/code&gt; - Find where the query beginning matches&lt;/li&gt;
&lt;li&gt;&lt;code&gt;alignEnd(querySuffix, target)&lt;/code&gt; - Find where the query end matches (by reversing and aligning)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;findMatchRegion(query, target)&lt;/code&gt; - Combine both to get the full match region&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The algorithm runs in O(mn) time where m and n are the lengths of the strings being aligned. For typical text selections (tens to hundreds of characters), this is instantaneous. And unlike the normalization approach, it's robust and correct!&lt;/p&gt;
</content></entry><entry><title>Canvas Chat: A Visual Interface for Thinking with LLMs</title><link href="https://ericmjl.github.io/blog/2025/12/31/canvas-chat-a-visual-interface-for-thinking-with-llms/" rel="alternate"/><updated>2025-12-31T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:d2aa2631-1832-38cd-8942-aab5690ab5ca</id><content type="html">&lt;p&gt;I've been mulling over this idea since last year January: A visual, nonlinear interface for LLM conversations—something like an infinite canvas where you could branch, merge, and see the shape of your thinking. It stayed in the "someday" pile because the implementation cost felt too high for a speculative side project; I wasn't skilled in browser technologies or anything UI-related.&lt;/p&gt;
&lt;p&gt;Then came the Christmas break ultralearning exercise I documented in &lt;a href="https://ericmjl.github.io/blog/2025/12/28/you-can-just-make-stuff-with-opencode-and-claude-opus-4-5/"&gt;my recent blog post about building with OpenCode and Claude Opus 4.5&lt;/a&gt;. Pressure-testing Opus 4.5 made me realize it was finally feasible to spend a day trying to make this work. I pushed Canvas Chat from idea to working prototype in about 24 hours of actual building time, and &lt;a href="https://ericmjl--canvas-chat-fastapi-app.modal.run/"&gt;then another 24 hrs to get it up on Modal&lt;/a&gt; and add in many, many refinements, each of which may have taken me multiple weeks. The final result is this:&lt;/p&gt;
&lt;p&gt;&lt;img src="canvas-chat-overview.webp" alt="Canvas Chat overview showing nodes connected in a directed graph"&gt;&lt;/p&gt;
&lt;p&gt;But before I explain what I built, let me explain &lt;em&gt;why&lt;/em&gt; I wanted it in the first place.&lt;/p&gt;
&lt;h2 id="the-job-to-be-done"&gt;The job to be done&lt;/h2&gt;&lt;p&gt;Clayton Christensen's &lt;a href="https://hbr.org/2016/09/know-your-customers-jobs-to-be-done"&gt;Jobs to Be Done&lt;/a&gt; framework asks: what job is the customer hiring this product to do? For Canvas Chat, the job isn't "chat with an LLM"—ChatGPT already does that fine. The job is: &lt;strong&gt;think through a complex problem where the exploration is nonlinear.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Here's the struggling moment. You're deep in a conversation with Claude or GPT, and you want to try a different framing of your question. But if you do, you'll lose the current thread. Or an LLM gives you a list of ten ideas, one catches your eye, and you want to drill into it—but the conversation keeps scrolling and you lose the overview. Or you've been exploring a problem across three separate chat sessions and now you need to synthesize, but you can't see them together.&lt;/p&gt;
&lt;p&gt;Linear chat actively works against this kind of thinking. It forces linear structure onto nonlinear exploration. You end up managing context in your head, copy-pasting between windows, losing track of which threads went where.&lt;/p&gt;
&lt;p&gt;Canvas Chat exists to solve that. When your thinking branches in multiple directions, it keeps all the threads visible and connected so you don't lose context and can synthesize across them.&lt;/p&gt;
&lt;h2 id="how-it-works"&gt;How it works&lt;/h2&gt;&lt;p&gt;Canvas Chat is an infinite canvas where conversations are nodes in a directed graph. You type a message, it appears as a node. The LLM's response appears as another node, connected by an edge. So far, standard. But then:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Branch from any node.&lt;/strong&gt; Click reply on any message, and your new message connects to that point, not the end of the conversation. The response branches off visually. Try two different prompts from the same starting point and see both branches side by side.&lt;/p&gt;
&lt;p&gt;&lt;img src="branch-from-node.webp" alt="Branching from a node to create parallel conversation threads"&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Highlight and branch.&lt;/strong&gt; Select text within a node, and a tooltip appears. Type a follow-up question, and Canvas Chat creates a highlight node (showing the excerpt with a blockquote) plus your question, plus the LLM response. The original node stays intact. This works especially well when an LLM gives a list of ideas and you want to drill into one without losing the overview.&lt;/p&gt;
&lt;p&gt;&lt;img src="highlight-and-branch-tooltip.webp" alt="Tooltip appearing when text is selected within a node"&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="highlight-and-branch-result.webp" alt="Result of highlight and branch showing the blockquote excerpt"&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multi-select for merge context.&lt;/strong&gt; Cmd-click multiple nodes, then type. The new message connects to all selected nodes, and the LLM sees the full ancestry of every selected node. I use this to synthesize: select two branches that went in different directions, ask "What do these approaches have in common?" The context includes everything that led to both.&lt;/p&gt;
&lt;p&gt;&lt;img src="multi-select-merge.webp" alt="Multiple nodes selected for merge context synthesis"&gt;&lt;/p&gt;
&lt;h2 id="context-flows-through-the-graph"&gt;Context flows through the graph&lt;/h2&gt;&lt;p&gt;When you send a message, Canvas Chat walks the DAG backward from your selected node(s), collecting all ancestors. It sorts them by creation time and sends them to the LLM as conversation history. If you've selected multiple nodes (a merge), the context is the union of all their ancestors, deduplicated.&lt;/p&gt;
&lt;p&gt;The practical effect: the LLM always knows how you arrived at the current question, even if the path is nonlinear. Branch from a discussion about protein folding dynamics, ask a follow-up about computational costs, and the context includes the protein folding discussion. No manual copy-paste.&lt;/p&gt;
&lt;h2 id="matrix-evaluation"&gt;Matrix evaluation&lt;/h2&gt;&lt;p&gt;This feature came out of a specific struggling moment: evaluating many options against many criteria and losing track of which combinations I'd thought through.&lt;/p&gt;
&lt;p&gt;Select one or more nodes as context, type &lt;code&gt;/matrix &amp;lt;and then put additional instructions you're looking to fill out here&amp;gt;&lt;/code&gt;. Canvas Chat parses out the list items and shows a confirmation modal where you can remove items or swap rows/columns. Click create, and a matrix node appears.&lt;/p&gt;
&lt;p&gt;&lt;img src="matrix-evaluation-modal.webp" alt="Matrix evaluation modal for configuring rows and columns"&gt;&lt;/p&gt;
&lt;p&gt;Each cell has a "+" button. Click it and the LLM fills that cell, seeing the matrix context you provided, the row item, the column item, and the full DAG history from the source nodes. "Fill All" processes every empty cell sequentially.&lt;/p&gt;
&lt;p&gt;Click any filled cell to see the full text. "Pin to Canvas" extracts that evaluation into a standalone node, which you can then branch from. Say you're comparing business ideas against criteria, one cell says "strong market fit with enterprise customers," you want to dig into that—pin and branch.&lt;/p&gt;
&lt;p&gt;&lt;img src="matrix-evaluation-filled.webp" alt="Matrix evaluation with cells filled by the LLM"&gt;&lt;/p&gt;
&lt;h2 id="web-search-and-deep-research"&gt;Web search and deep research&lt;/h2&gt;&lt;p&gt;Canvas Chat integrates Exa's APIs for two slash commands:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;/search &amp;lt;query&amp;gt;&lt;/code&gt; runs a neural search and creates a Search node with the query, plus Reference nodes for each result. Click "Fetch &amp;amp; Summarize" on any reference to grab the full page content and summarize it.&lt;/p&gt;
&lt;p&gt;&lt;img src="web-search-results.webp" alt="Web search results showing reference nodes"&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;/research &amp;lt;topic&amp;gt;&lt;/code&gt; kicks off Exa's Research API, which performs multi-step research with multiple queries. The results stream into a Research node with inline source citations.&lt;/p&gt;
&lt;p&gt;&lt;img src="deep-research-node.webp" alt="Deep research node with inline citations"&gt;&lt;/p&gt;
&lt;p&gt;If you have nodes selected when you run these commands, Canvas Chat uses an LLM to refine your query using the selected text as context. Highlight "CCNOT gate" and type &lt;code&gt;/search how does this work&lt;/code&gt;, and it rewrites the query to "how Toffoli gate CCNOT quantum computing works" before searching.&lt;/p&gt;
&lt;h2 id="local-first-and-multi-provider"&gt;Local-first and multi-provider&lt;/h2&gt;&lt;p&gt;&lt;img src="settings-api-keys.webp" alt="Settings panel showing API key configuration"&gt;&lt;/p&gt;
&lt;p&gt;All session data lives in IndexedDB. No server-side storage, no accounts. Export sessions as &lt;code&gt;.canvaschat&lt;/code&gt; JSON files. API keys live in localStorage and are sent with each request.&lt;/p&gt;
&lt;p&gt;The server is stateless: it proxies LLM calls via LiteLLM and handles the Exa integration, but never stores conversation data. You can deploy it yourself on Modal with a single command.&lt;/p&gt;
&lt;p&gt;Canvas Chat dynamically fetches available models from each provider when you enter an API key. OpenAI, Anthropic, Google (Gemini), Groq, GitHub Models, and local Ollama instances (when running on localhost) all work. Switch models mid-conversation to compare outputs.&lt;/p&gt;
&lt;h2 id="what-building-this-taught-me"&gt;What building this taught me&lt;/h2&gt;&lt;p&gt;This project reinforced something I wrote about in &lt;a href="../../28/you-can-just-make-stuff-with-opencode-and-claude-opus-4-5/"&gt;the "I don't code anymore, I build" post&lt;/a&gt;: I stayed in product builder brain throughout. I didn't have strong opinions about whether the JavaScript was idiomatic because I don't know what idiomatic JavaScript looks like. I just knew whether the feature worked.&lt;/p&gt;
&lt;p&gt;When something broke, I'd describe the symptoms and let Opus 4.5 debug in as much detail as I can manage. When I wanted a new interaction pattern, I'd describe what it should feel like and watch it materialize. The creative work — deciding what nonlinear chat should &lt;em&gt;be&lt;/em&gt; — remained human. The mechanical translation got delegated.&lt;/p&gt;
&lt;p&gt;Canvas Chat is the kind of project I wouldn't have attempted before because the implementation cost exceeded the payoff. Now it didn't.&lt;/p&gt;
&lt;h2 id="try-it"&gt;Try it&lt;/h2&gt;&lt;p&gt;Canvas Chat is open source. Run it locally:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;git&lt;span class="w"&gt; &lt;/span&gt;clone&lt;span class="w"&gt; &lt;/span&gt;https://github.com/ericmjl/canvas-chat.git
&lt;span class="nb"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;canvas-chat
pixi&lt;span class="w"&gt; &lt;/span&gt;run&lt;span class="w"&gt; &lt;/span&gt;dev
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Add your API keys in settings and go. The deployed version runs on &lt;a href="https://ericmjl--canvas-chat-fastapi-app.modal.run/"&gt;Modal&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If you try it, I want to hear what works and what doesn't! You can get in touch with me via &lt;a href="https://ericmjl--shortmail-run-app.modal.run/send/cce87ae9c1d7"&gt;Shortmail&lt;/a&gt;, or file an issue on the &lt;a href="https://github.com/ericmjl/canvas-chat"&gt;Github repo&lt;/a&gt;.&lt;/p&gt;
</content></entry><entry><title>You Can Just Make Stuff with OpenCode and Claude Opus 4.5</title><link href="https://ericmjl.github.io/blog/2025/12/28/you-can-just-make-stuff-with-opencode-and-claude-opus-4-5/" rel="alternate"/><updated>2025-12-28T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:c6c818be-320a-3218-b326-65fc288e1a41</id><content type="html">&lt;p&gt;&lt;a href="https://www.linkedin.com/feed/update/urn:li:activity:7408389557915799552?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A7408389557915799552%2C7410474533469597697%29&amp;amp;dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287410474533469597697%2Curn%3Ali%3Aactivity%3A7408389557915799552%29"&gt;Tommy Tang asked me&lt;/a&gt; about my opinions on OpenCode, so here's what I've learned after spending significant time with &lt;a href="https://opencode.ai/"&gt;OpenCode&lt;/a&gt; and Claude Opus 4.5.&lt;/p&gt;
&lt;h2 id="i-don-t-code-anymore-i-build"&gt;I don't code anymore, I build&lt;/h2&gt;&lt;p&gt;This is the punchline, so let me start with it. I've shifted from writing code to directing its creation. The change happened gradually, then all at once. I used to think about syntax, edge cases, and implementation details. Now I think about what I want to exist, describe it clearly, and watch it materialize.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.biblegateway.com/passage/?search=Genesis%201&amp;amp;version=NIV"&gt;Genesis 1:3&lt;/a&gt; describes this pattern at a cosmic scale:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;"And God said, 'Let there be light,' and there was light."&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Working with Claude Opus 4.5 through OpenCode feels like a microcosm of that creative act.&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;Eric said, "Let there be a feature," and there was the feature, in code.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I'm not claiming divinity here, just noting that the creative pattern of speaking things into existence has become surprisingly literal in my daily work.&lt;/p&gt;
&lt;h2 id="the-tools-opencode-and-claude-opus-4.5"&gt;The tools: OpenCode and Claude Opus 4.5&lt;/h2&gt;&lt;p&gt;Like &lt;a href="https://www.youtube.com/watch?v=4AyM_3SK31w&amp;amp;t=1263s"&gt;Theo Brown from t3.gg&lt;/a&gt;, I've settled on Claude Opus 4.5 as my primary model for coding tasks. It just knows what to do. I've stopped trying to micro-manage the model's actions because it handles most tasks autonomously and correctly. When I ask for a refactor, it refactors. When I describe a feature, it implements it. The gap between intention and execution has shrunk to almost nothing.&lt;/p&gt;
&lt;p&gt;Other models require more hand-holding. Opus 4.5 seems to have internalized enough software engineering patterns that I can trust it to make reasonable architectural decisions without constant course corrections. I can literally ask it to "do the docs, keep things up-to-date, and also give me a document that has an overview of code organization and architecture." It just goes to town autonomously. No step-by-step prompting, no breaking the task into smaller pieces. I describe the outcome I want and it figures out the path.&lt;/p&gt;
&lt;p&gt;The tooling layer matters too. &lt;a href="https://opencode.ai/"&gt;OpenCode&lt;/a&gt; orchestrates the AI coding in a way that feels natural. The tools it calls are always logical, the reasoning traces are transparent, and the execution flow makes sense. It shows a running list of modified files, giving me context about what's changing without running &lt;code&gt;git status&lt;/code&gt; constantly. Context compaction lets me stay in one long-running session without hitting token limits. I've thrown out the old playbook of "switch sessions when you approach the context window." Now I only switch when I want to do something entirely different.&lt;/p&gt;
&lt;p&gt;My setup: OpenCode with auto-updating, GitHub Copilot Pro as the LLM provider (routing to Opus 4.5), running inside a &lt;a href="https://github.com/tmux/tmux"&gt;tmux&lt;/a&gt; session for persistence. Each repo gets an AGENTS.md file where I encode my preferences and patterns - the model's training data for my specific context. Opus 4.5 actually respects what's in there, unlike some other models that seem to ignore custom instructions.&lt;/p&gt;
&lt;h2 id="ten-days-of-deliberate-practice"&gt;Ten days of deliberate practice&lt;/h2&gt;&lt;p&gt;I decided to pressure-test the "I build" claim over the holidays. Ten days, December 19-28, using OpenCode as my primary development interface. The goal: see how much I could actually ship.&lt;/p&gt;
&lt;p&gt;The answer surprised me. Across six repositories, I pushed over 150 commits spanning infrastructure work, documentation, greenfield apps, and maintenance. Here's what emerged:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A &lt;a href="https://ericmjl.github.io/2025-12-ski-trip-website/"&gt;ski trip coordination website&lt;/a&gt;&lt;/strong&gt; (59 commits). My family was heading to New Hampshire for a week. Normally I'd have used a shared Google Doc for the itinerary. Instead, I built a full website with recipe modals, restaurant links with Apple and Google Maps integration, a photo album with lightbox navigation, automatic thumbnail generation, and a hero video background. I updated it live during the trip - adding photos, adjusting the grocery list, swapping menu items. The implementation cost would have been absurd for a week-long trip before. Now the jazz and snazz was well worth the effort - my family actually enjoyed using it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A teaching clock app for my kids&lt;/strong&gt; (2 commits, but a complete app). An &lt;a href="https://ericmjl.github.io/teaching-clock/"&gt;analog clock trainer&lt;/a&gt; plus a &lt;a href="https://ericmjl.github.io/teaching-clock/puzzle.html"&gt;jigsaw puzzle game&lt;/a&gt; with difficulty levels and themes. Pure JavaScript and CSS - exactly the kind of project my decade-old "no JavaScript" rule would have blocked. The model wrote it; I directed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/pyjanitor-devs/pyjanitor"&gt;pyjanitor&lt;/a&gt; infrastructure&lt;/strong&gt; (40 commits). Currency symbol support for international formats. Automated patch releases on every merge. Test isolation fixes. And a major expansion of AGENTS.md into what I now think of as the repository's "agent constitution" - a document that tells AI assistants how to work within this specific codebase.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A new &lt;a href="https://github.com/conda-forge/staged-recipes/pull/31776"&gt;conda-forge package&lt;/a&gt;&lt;/strong&gt; for janitor-rs. The model handled the unfamiliar territory of Rust packaging and conda-forge recipe formats. I was the novice here; it was the guide. This role reversal keeps happening - when I set up PostHog analytics or migrated to GA4 on my website, the model walked me through each step, explained what I was doing and why, and waited for confirmation before proceeding. The expert-novice relationship flips depending on who knows more about the task at hand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A &lt;a href="../../27/how-i-themed-my-tmux-with-opencode-and-claude/"&gt;custom tmux status bar&lt;/a&gt;&lt;/strong&gt; with Nord colors, powerline arrows, and smooth color transitions. Pure aesthetic indulgence - the kind of project I'd never have prioritized before because the implementation cost exceeded the payoff. Now it didn't.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/ericmjl/canvas-chat"&gt;Canvas Chat&lt;/a&gt;&lt;/strong&gt; (13 commits in 24 hours). A visual non-linear chat interface - think infinite canvas meets LLM conversation. Resizable nodes, trackpad gestures, streaming responses, web search via Exa, session management. FastAPI backend, vanilla JS frontend. Another "no JavaScript" rule violation, and another project that went from idea to working prototype in a single day.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Smaller fixes across &lt;a href="https://github.com/ericmjl/llamabot"&gt;llamabot&lt;/a&gt;&lt;/strong&gt; (better error messages) and &lt;strong&gt;&lt;a href="https://github.com/ericmjl/website"&gt;my website&lt;/a&gt;&lt;/strong&gt; (PostHog analytics, GA4 migration, blog posts).&lt;/p&gt;
&lt;p&gt;The variety matters. This wasn't one type of project where I got lucky. It was infrastructure, documentation, greenfield consumer apps, packaging for an ecosystem I rarely touch, and routine maintenance. The "I build" claim held up across all of them.&lt;/p&gt;
&lt;h2 id="from-engineer-brain-to-product-builder-brain"&gt;From engineer brain to product builder brain&lt;/h2&gt;&lt;p&gt;Something shifted in how I think about these projects. Previously, I'd worry about &lt;em&gt;how&lt;/em&gt; a thing was built - the engineer brain obsessing over implementation details, code structure, idiomatic patterns. Now I've switched to &lt;em&gt;what&lt;/em&gt; was built, &lt;em&gt;why&lt;/em&gt; I want it built, and &lt;em&gt;does it get the job done&lt;/em&gt; - the product builder's brain.&lt;/p&gt;
&lt;p&gt;This is especially true for the ski website and Canvas Chat, both built with web technologies (HTML, JS, CSS) that I'm not deeply familiar with. Ironically, my unfamiliarity frees me from micro-managing the implementation. I don't have strong opinions about whether the JavaScript is idiomatic because I don't know what idiomatic JavaScript looks like. I just know whether the feature works.&lt;/p&gt;
&lt;p&gt;But there's a latent risk here. The code might not follow best practices - lots of duplication, poor separation of concerns, missing edge cases. So I fall back on &lt;em&gt;principles&lt;/em&gt; I picked up from years of Python: refactoring, documentation, testing. I stay at that level of nudging Opus 4.5 - "look for places to refactor," "document this module," "add tests for this functionality" - but I stay out of the nitty-gritty implementation. The principles transfer even when the language doesn't.&lt;/p&gt;
&lt;h2 id="how-my-review-process-changed"&gt;How my review process changed&lt;/h2&gt;&lt;p&gt;Here's something I didn't expect: I don't scrutinize the code as tightly as I used to during active development. Instead, I read the reasoning traces first. The model's chain of thought tells me whether my codebase is heading in the right direction. If the reasoning is coherent and addresses the right concerns, the code will reflect what I want. If the reasoning seems confused or takes weird detours, something's wrong and I need to dig deeper.&lt;/p&gt;
&lt;p&gt;This inverts the traditional development loop. I used to read code to understand what the computer would do. Now I read reasoning to understand what the model understood and decided. The code review happens afterward, and it's lighter because the reasoning already told me whether we're on track.&lt;/p&gt;
&lt;p&gt;When I want to catch issues that slipped through, I start a fresh session. A new context window acts like a fresh pair of eyes - the model hasn't been primed by the conversation that led to the current implementation, so it can spot inconsistencies that were invisible during the creative flow. This parallels the old advice about stepping away from code before reviewing it, except now the "stepping away" happens by instantiating a new session rather than waiting for my own brain to reset.&lt;/p&gt;
&lt;h2 id="unlearning-old-assumptions"&gt;Unlearning old assumptions&lt;/h2&gt;&lt;p&gt;&lt;a href="https://x.com/bcherny/status/2004626064187031831"&gt;Boris Cherny recently had a Twitter exchange with Andrej Karpathy&lt;/a&gt; that resonated with me. Boris observed that newer coworkers and even new grads who don't make assumptions about what the model can and can't do are often able to use it most effectively. They don't carry "legacy memories formed when using old models." Every month or two, models get better, and those of us who've been using them longest have to actively unlearn outdated limitations.&lt;/p&gt;
&lt;p&gt;I've caught myself doing this repeatedly. Back in grad school around 2015, I tried building &lt;a href="https://d3js.org/"&gt;d3.js&lt;/a&gt; visualizations and struggled to adjust to JavaScript's syntax coming from Python. I decided to focus on getting better at Python first and gave myself a "no JavaScript" rule wherever possible. That constraint made sense at the time. It makes no sense now. The model writes JavaScript just fine. My decade-old "no JavaScript" policy was a legacy memory holding me back from building things that would actually benefit from running in the browser.&lt;/p&gt;
&lt;p&gt;The mental work of re-adjusting expectations is real. I have to keep asking myself: would I have avoided this six months ago because the model couldn't handle it, or because I assumed it couldn't? The answer is increasingly the latter.&lt;/p&gt;
&lt;p&gt;There's a flip side to this unlearning, though. Working in JavaScript land forced me to learn the language of the web to achieve the same precision and fluency I have with Python. I found myself picking up patterns I'd avoided for years: the browser console for debugging, DOM element manipulation, CSS transitions I didn't know existed, the JS package ecosystem. The model writes the code, but I still need enough vocabulary to direct it well and recognize when something's off. Unlearning old constraints doesn't mean staying ignorant of new territory - it means finally having a reason to explore it.&lt;/p&gt;
&lt;h2 id="what-this-means"&gt;What this means&lt;/h2&gt;&lt;p&gt;The shift from "I code" to "I build" isn't just semantic. It reflects a genuine change in what I spend my attention on. Less time on syntax and implementation details. More time on architecture, requirements, and verification. The creative work remains human. The mechanical translation has been delegated.&lt;/p&gt;
&lt;p&gt;I'm still learning how to use this effectively. But the trajectory is clear: the gap between imagining software and having software continues to shrink.&lt;/p&gt;
</content></entry><entry><title>How I Themed My tmux with OpenCode + Claude (And When to Switch Models)</title><link href="https://ericmjl.github.io/blog/2025/12/27/how-i-themed-my-tmux-with-opencode-and-claude/" rel="alternate"/><updated>2025-12-27T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:9ed7b606-5f8b-3aef-8718-ba121e610e6e</id><content type="html">&lt;p&gt;I had a beautiful tmux status bar on my old laptop. Nord colors, powerline arrows, clean and minimal. The kind that makes you feel like a proper terminal power user.&lt;/p&gt;
&lt;p&gt;When I got a new machine back in April, I was too lazy to set up tmux properly. The sensible thing would have been to spend five minutes copying over my old config. Instead, eight months later, I finally spent an hour pair-programming the whole thing from scratch with &lt;a href="https://github.com/sst/opencode"&gt;OpenCode&lt;/a&gt; and Claude.&lt;/p&gt;
&lt;p&gt;Why? Honestly, I wanted to try out a new tool. The irony isn't lost on me.&lt;/p&gt;
&lt;h2 id="the-setup"&gt;The Setup&lt;/h2&gt;&lt;p&gt;OpenCode is a CLI tool that lets you interact with Claude directly from your terminal. Perfect for this kind of task: I'm already in the terminal configuring tmux, so having my AI pair programmer right there keeps the feedback loop tight. Describe what I want. See the change. Describe what's wrong, with precision. Iterate. No context switching to a browser.&lt;/p&gt;
&lt;p&gt;That tight loop is what let me stay in the creative headspace. I could say things like "I want the arrows to overlap like in this screenshot" or "the colors feel too muted, try the frost blue from Nord" without knowing the exact syntax. Claude translated my aesthetic intent into working config.&lt;/p&gt;
&lt;p&gt;The other superpower: model switching. OpenCode lets you flip between any models you have API keys for. For this session, I toggled between Claude Sonnet (fast, good for quick iterations) and Claude Opus (slower, but sharper for complex debugging). This turned out to be crucial.&lt;/p&gt;
&lt;h2 id="starting-with-research"&gt;Starting with Research&lt;/h2&gt;&lt;p&gt;First, I asked Sonnet to search online for tmux status bar customization. It pulled resources from the official tmux wiki and various tutorials, giving me a foundation: &lt;code&gt;status-left&lt;/code&gt;, &lt;code&gt;status-right&lt;/code&gt;, &lt;code&gt;window-status-format&lt;/code&gt;, color options, the basics.&lt;/p&gt;
&lt;p&gt;Armed with that, we dove in.&lt;/p&gt;
&lt;h2 id="first-attempt-with-a-custom-theme"&gt;First attempt with a custom theme&lt;/h2&gt;&lt;p&gt;Claude created a custom dark theme inspired by Catppuccin colors. Worked immediately:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;status-style&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;bg=#1e1e2e,fg=#cdd6f4&amp;quot;&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;status-left&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;#[fg=#89b4fa,bold] #S #[fg=#a6e3a1]@ #H&amp;quot;&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;status-right&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;#[fg=#f9e2af]%a %b %d #[fg=#89b4fa]%H:%M&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Clean. Functional. Pretty. But I wanted more: those beautiful powerline arrows flowing between segments. That's when things got interesting.&lt;/p&gt;
&lt;h2 id="the-powerline-saga"&gt;The Powerline Saga&lt;/h2&gt;&lt;p&gt;Claude suggested &lt;code&gt;powerline-go&lt;/code&gt;, a Go-based powerline prompt generator. We installed it via Homebrew (not pip, since &lt;a href="https://ericmjl.github.io/blog/2024/8/16/its-time-to-try-out-pixi/"&gt;I keep my system Python-free&lt;/a&gt;):&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;brew&lt;span class="w"&gt; &lt;/span&gt;install&lt;span class="w"&gt; &lt;/span&gt;powerline-go
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Updated the tmux config to call powerline-go for the status bar. Reloaded. And... disaster. Instead of beautiful arrows, raw escape codes:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;[38;5;15m[48;5;4m ericmjl [38;5;4m[48;5;0m...
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The terminal was spitting out ANSI codes instead of interpreting them. We tried various fixes, but powerline-go simply wasn't designed for tmux status bars; it's meant for shell prompts. Back to square one.&lt;/p&gt;
&lt;h2 id="trying-the-tmux-powerline-plugin"&gt;Trying the tmux-powerline Plugin&lt;/h2&gt;&lt;p&gt;Next attempt: the actual &lt;code&gt;tmux-powerline&lt;/code&gt; plugin via TPM (Tmux Plugin Manager):&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;git&lt;span class="w"&gt; &lt;/span&gt;clone&lt;span class="w"&gt; &lt;/span&gt;https://github.com/tmux-plugins/tpm&lt;span class="w"&gt; &lt;/span&gt;~/.tmux/plugins/tpm
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Added the plugin, pressed &lt;code&gt;C-b I&lt;/code&gt; to install, and... the status bar exploded with information. IP addresses, weather, load averages, hostname. Way too much. I asked Claude to simplify and switch to Nord colors.&lt;/p&gt;
&lt;p&gt;We created a custom theme at &lt;code&gt;~/.config/tmux-powerline/themes/nord.sh&lt;/code&gt;, updated the config, reloaded tmux. Nothing changed. The theme wasn't loading. Killed the server entirely. Restarted. Still the old crowded theme.&lt;/p&gt;
&lt;p&gt;This is where Sonnet started struggling. Same fixes over and over: reload the config, check the theme path, restart tmux. Loop after loop of suggestions that weren't working.&lt;/p&gt;
&lt;h2 id="the-model-switch-from-sonnet-to-opus"&gt;The model switch from Sonnet to Opus&lt;/h2&gt;&lt;p&gt;I noticed Sonnet spinning its wheels. Same suggestions, same non-results. Time to switch.&lt;/p&gt;
&lt;p&gt;The difference was immediate. Instead of repeating failed approaches, Opus stepped back and proposed something different entirely: ditch the plugin and go native. Tmux's built-in formatting is powerful enough to create powerline-style status bars without any plugins. We just needed the right Unicode characters and color transitions.&lt;/p&gt;
&lt;p&gt;This stuck with me: Sonnet is fantastic for speed and quick iterations, but when you're stuck in a loop, Opus brings the lateral thinking to break out.&lt;/p&gt;
&lt;h2 id="going-native-as-the-winning-approach"&gt;Going native as the winning approach&lt;/h2&gt;&lt;p&gt;Fresh start. Clean native tmux config. The key insight was understanding how powerline arrows actually work: the arrow character's foreground color matches the background of the segment it's coming from, and its background matches what it's going into.&lt;/p&gt;
&lt;p&gt;Here's the final status-left (session name with powerline arrow):&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;status-left&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;#[fg=#2e3440,bg=#5e81ac,bold]  #S #[fg=#5e81ac,bg=#2e3440]\ue0b0&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The window formats, with arrows on both sides so they flow into neighboring elements:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# Inactive windows&lt;/span&gt;
setw&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;window-status-format&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;#[fg=#2e3440,bg=#3b4252]\ue0b0#[fg=#d8dee9,bg=#3b4252] #I #W #[fg=#3b4252,bg=#2e3440]\ue0b0&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# Active window (cyan highlight)&lt;/span&gt;
setw&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;window-status-current-format&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;#[fg=#2e3440,bg=#88c0d0]\ue0b0#[fg=#2e3440,bg=#88c0d0,bold] #I #W #[fg=#88c0d0,bg=#2e3440]\ue0b0&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And the right side (battery, date, time) using left-pointing arrows and a smooth Nord color gradient:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-g&lt;span class="w"&gt; &lt;/span&gt;status-right&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;#[fg=#a3be8c,bg=#2e3440]\ue0b2#[fg=#2e3440,bg=#a3be8c,bold] 󰁹 #(pmset -g batt | grep -o &amp;#39;[0-9]*%%&amp;#39; | head -1) #[fg=#5e81ac,bg=#a3be8c]\ue0b2#[fg=#d8dee9,bg=#5e81ac] %b %d #[fg=#88c0d0,bg=#5e81ac]\ue0b2#[fg=#2e3440,bg=#88c0d0,bold] %H:%M &amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="the-final-result"&gt;The Final Result&lt;/h2&gt;&lt;p&gt;After all that iteration, here's what my tmux status bar looks like:&lt;/p&gt;
&lt;div style="background: #2e3440; font-family: 'JetBrains Mono', 'Fira Code', monospace; padding: 8px 0; display: flex; flex-wrap: nowrap; justify-content: space-between; align-items: center; border-radius: 4px; overflow: hidden; min-width: 0;"&gt;
  &lt;div style="display: flex; flex-wrap: nowrap; align-items: center; height: 24px; flex-shrink: 0;"&gt;
    &lt;span style="background: #5e81ac; color: #2e3440; padding: 0 12px; font-weight: bold; height: 100%; display: flex; align-items: center; white-space: nowrap;"&gt;system-config&lt;/span&gt;
    &lt;div style="width: 0; height: 0; border-top: 12px solid transparent; border-bottom: 12px solid transparent; border-left: 12px solid #5e81ac; flex-shrink: 0; position: relative; z-index: 2;"&gt;&lt;/div&gt;
    &lt;span style="background: #88c0d0; color: #2e3440; padding: 0 12px; font-weight: bold; height: 100%; display: flex; align-items: center; white-space: nowrap; margin-left: -12px; padding-left: 20px;"&gt;1 opencode&lt;/span&gt;
    &lt;div style="width: 0; height: 0; border-top: 12px solid transparent; border-bottom: 12px solid transparent; border-left: 12px solid #88c0d0; flex-shrink: 0;"&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div style="display: flex; flex-wrap: nowrap; align-items: center; height: 24px; flex-shrink: 0;"&gt;
    &lt;div style="width: 0; height: 0; border-top: 12px solid transparent; border-bottom: 12px solid transparent; border-right: 12px solid #a3be8c; flex-shrink: 0;"&gt;&lt;/div&gt;
    &lt;span style="background: #a3be8c; color: #2e3440; padding: 0 12px; font-weight: bold; height: 100%; display: flex; align-items: center; white-space: nowrap; padding-right: 20px;"&gt;🔋 100%&lt;/span&gt;
    &lt;div style="width: 0; height: 0; border-top: 12px solid transparent; border-bottom: 12px solid transparent; border-right: 12px solid #5e81ac; flex-shrink: 0; margin-left: -12px; position: relative; z-index: 2;"&gt;&lt;/div&gt;
    &lt;span style="background: #5e81ac; color: #d8dee9; padding: 0 12px; height: 100%; display: flex; align-items: center; white-space: nowrap; padding-right: 20px;"&gt;Dec 23&lt;/span&gt;
    &lt;div style="width: 0; height: 0; border-top: 12px solid transparent; border-bottom: 12px solid transparent; border-right: 12px solid #88c0d0; flex-shrink: 0; margin-left: -12px; position: relative; z-index: 2;"&gt;&lt;/div&gt;
    &lt;span style="background: #88c0d0; color: #2e3440; padding: 0 12px; font-weight: bold; height: 100%; display: flex; align-items: center; white-space: nowrap;"&gt;06:05&lt;/span&gt;
  &lt;/div&gt;
&lt;/div&gt;&lt;p&gt;Session name in frost blue on the left. Active window in cyan. Right side flows through battery (green), date (blue), and time (cyan). All connected by powerline arrows with smooth color transitions. &lt;em&gt;(I asked Claude to recreate the status bar in HTML so I wouldn't have to screenshot it for the blog.)&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="what-i-took-away"&gt;What I Took Away&lt;/h2&gt;&lt;p&gt;There's a growing conversation about AI-assisted programming: the tight feedback loops, model selection strategies, iterative workflows. I've written about some of these patterns myself. But this session crystallized something different.&lt;/p&gt;
&lt;p&gt;I can express my creativity on a computer screen more easily than ever before.&lt;/p&gt;
&lt;p&gt;I'm not a designer. CSS is foreign to me, hex color codes don't stick in my head, and tmux's formatting syntax is arcane. But I have taste. I know what looks good. Years of admiring beautiful terminals gave me a mental mood board. What I lacked was the technical fluency to make it real.&lt;/p&gt;
&lt;p&gt;AI bridged that gap. Throughout this session I worked like a designer: describing aesthetics, pointing at visual problems, directing iteration. "The arrows should overlap." "That cyan is too bright." "Make the battery segment green." Claude handled implementation. I stayed in the creative headspace.&lt;/p&gt;
&lt;p&gt;Iteration surfaces what you actually want.&lt;/p&gt;
&lt;p&gt;This surprised me. I didn't start with a complete vision, just a vague sense of "Nord colors, powerline arrows, clean and minimal." But each rapid cycle surfaced preferences I didn't know I had. The arrows need to overlap. The active window should pop more. The right side needs a color gradient. None of these were requirements I could have articulated upfront. They emerged through seeing and reacting.&lt;/p&gt;
&lt;p&gt;Bits and bytes have never been cheaper to produce. AI can generate config files, CSS, code, whatever. But aesthetics and judgment? Those remain expensive. The scarce resource isn't the implementation anymore. It's knowing what you want and recognizing when you've found it.&lt;/p&gt;
&lt;p&gt;AI doesn't replace that judgment. It amplifies it by removing the implementation friction that used to slow the creative loop down.&lt;/p&gt;
&lt;p&gt;The whole session took about an hour, failed attempts included. Without AI pair programming, I'd probably still be reading documentation. Instead, I have a beautiful terminal, and a new appreciation for what becomes possible when the gap between creative vision and technical implementation shrinks to nearly nothing.&lt;/p&gt;
</content></entry><entry><title>Two years of weekly blogging and what 2025 taught me</title><link href="https://ericmjl.github.io/blog/2025/12/25/two-years-of-weekly-blogging-and-what-2025-taught-me/" rel="alternate"/><updated>2025-12-25T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:af479e56-5fbc-334a-b7c8-6532edd8477b</id><content type="html">&lt;p&gt;Last year, I challenged myself to write one blog post per week,
and I hit 53 posts by the end of 2024.
This year, I doubled down on that commitment
and wrote 50 posts in 2025.
Including this one, it's 51,
bringing me to 104 blog posts over two years.&lt;/p&gt;
&lt;h2 id="the-year-of-coding-agents"&gt;The year of coding agents&lt;/h2&gt;&lt;p&gt;Looking at my 2025 posts,
one theme dominates: &lt;strong&gt;coding agents&lt;/strong&gt;.
I wrote extensively about how to work with AI coding assistants,
from teaching them with AGENTS.md files
to letting them work autonomously.
This reflected a shift in how I work day-to-day.&lt;/p&gt;
&lt;p&gt;Some highlights from this theme:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/4/how-to-teach-your-coding-agent-with-agentsmd/"&gt;How to teach your coding agent with AGENTS.md&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/8/safe-ways-to-let-your-coding-agent-work-autonomously/"&gt;Safe ways to let your coding agent work autonomously&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/10/productive-patterns-for-agent-assisted-programming/"&gt;Productive Patterns for Agent-Assisted Programming&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/"&gt;How I Replaced 307 Lines of Agent Code with 4 Lines&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The shift from "AI as a tool" to "AI as a collaborator"
captures how my practice evolved this year.
I've gone from cautiously experimenting with Cursor
to having established patterns for multi-repository agent workflows.&lt;/p&gt;
&lt;h2 id="bayesian-methods-and-biological-applications"&gt;Bayesian methods and biological applications&lt;/h2&gt;&lt;p&gt;My work continued to inform my writing,
with several posts on applying Bayesian statistics to real lab problems.
The R2D2 prior posts were particularly satisfying to write
because I felt equipped with new theoretical knowledge that was directly applicable,
and I appreciated the mathematical aesthetics behind the approach:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/3/bayesian-superiority-estimation-with-r2d2-priors-a-practical-guide-for-protein-screening/"&gt;Bayesian Superiority Estimation with R2D2 Priors: A Practical Guide for Protein Screening&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/6/stop-guessing-at-priors-r2d2s-automated-approach-to-bayesian-modeling/"&gt;Stop guessing at priors: R2D2's automated approach to Bayesian modeling&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/5/from-data-chaos-to-statistical-clarity-a-laboratory-transformation-story/"&gt;From data chaos to statistical clarity: A laboratory transformation story&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I also explored the challenges of working with lab data,
including why preclinical experiments make ML challenging
and how to communicate effectively with lab scientists.&lt;/p&gt;
&lt;h2 id="tools-i-got-excited-about"&gt;Tools I got excited about&lt;/h2&gt;&lt;p&gt;Every year brings new tools that change how I work.
In 2025, two stood out.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://marimo.io/"&gt;Marimo&lt;/a&gt; is a reactive notebook tool
that I wrote about with enthusiasm,
and followed up with practical guidance on using coding agents to write Marimo notebooks.
The reactive execution model aligns well with how I think about data exploration.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://modal.com/"&gt;Modal&lt;/a&gt; is cloud computing that actually feels Pythonic.
My "Wow, Modal!" post captured the delight of finding infrastructure
that doesn't fight against my workflow.&lt;/p&gt;
&lt;h2 id="data-science-leadership-and-career"&gt;Data science leadership and career&lt;/h2&gt;&lt;p&gt;I continued writing about the human side of data science work,
including standardizing ways of working, communicating with lab scientists,
and navigating the biotech industry's ups and downs.
The year ended with
&lt;a href="https://ericmjl.github.io/blog/2025/12/17/the-selfish-reason-to-do-your-best-work/"&gt;The selfish reason to do your best work&lt;/a&gt;,
which synthesized lessons from a challenging year in biotech.&lt;/p&gt;
&lt;h2 id="looking-ahead-to-2026"&gt;Looking ahead to 2026&lt;/h2&gt;&lt;p&gt;After two years of writing almost weekly on whatever is on my mind,
I am adjusting my goals.
Next year, my attention shifts towards
(a) learning the fundamentals of quantum computing through an ultralearning project,
(b) writing more on data science leadership and career development to encourage colleagues navigating similar paths, and
(c) building out at least 10 experimental things with AI.
I am also dropping the goal of "one blog post per week" to four per month,
which brings me to a goal of 48 for 2026.
I am giving myself space to rest and strategically plan out writing going into 2026.&lt;/p&gt;
&lt;p&gt;Merry Christmas and a happy new year to all my readers!&lt;/p&gt;
&lt;h2 id="blog-posts-by-theme"&gt;Blog posts by theme&lt;/h2&gt;&lt;h3 id="biology-chemistry"&gt;Biology &amp;amp; Chemistry&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/4/what-makes-an-agent/"&gt;What makes an agent? (2025-01-04)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/19/why-data-from-preclinical-biotech-lab-experiments-make-machine-learning-challenging/"&gt;Why data from preclinical biotech lab experiments make machine learning challenging (2025-01-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/23/reliable-biological-data-requires-physical-quantities-not-statistical-artifacts/"&gt;Reliable biological data requires physical quantities, not statistical artifacts (2025-02-23)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/6/a-blueprint-for-data-driven-molecule-engineering/"&gt;A blueprint for data-driven molecule engineering (2025-03-06)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/3/bayesian-superiority-estimation-with-r2d2-priors-a-practical-guide-for-protein-screening/"&gt;Bayesian Superiority Estimation with R2D2 Priors: A Practical Guide for Protein Screening (2025-04-03)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/5/from-data-chaos-to-statistical-clarity-a-laboratory-transformation-story/"&gt;From data chaos to statistical clarity: A laboratory transformation story (2025-04-05)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/19/good-practices-for-ai-assisted-development-from-a-live-protein-calculator-demo/"&gt;Good practices for AI-assisted development from a live protein calculator demo (2025-04-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/27/build-your-own-tools/"&gt;Build your own tools! (2025-06-27)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/14/reflections-on-the-scipy-2025-conference/"&gt;Reflections on the SciPy 2025 Conference (2025-07-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/15/how-to-use-xarray-for-unified-laboratory-data-storage/"&gt;How to use xarray for unified laboratory data storage (2025-07-15)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/24/how-to-communicate-with-lab-scientists-when-youre-the-data-person/"&gt;How to communicate with lab scientists (when you're the data person) (2025-08-24)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/1/how-data-scientists-can-master-life-sciences-and-software-skills-for-biotech-using-ultralearning/"&gt;How data scientists can master life sciences and software skills for biotech using ultralearning (2025-10-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/2/what-does-it-take-to-build-a-statistics-agent/"&gt;What does it take to build a statistics agent? (2025-12-02)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="career-advice"&gt;Career Advice&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/13/writing-at-the-speed-of-thought/"&gt;Writing at the speed of thought (2025-01-13)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/17/why-you-should-take-part-in-the-scipy-sprints/"&gt;Why you should take part in the SciPy sprints! (2025-03-17)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/8/why-im-excited-for-scipy-2025/"&gt;Why I'm excited for SciPy 2025! (2025-05-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/14/reflections-on-the-scipy-2025-conference/"&gt;Reflections on the SciPy 2025 Conference (2025-07-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/15/data-scientists-arent-becoming-obsolete-in-the-llm-era/"&gt;Data scientists aren't becoming obsolete in the LLM era (2025-08-15)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/9/1/how-to-use-ai-to-accelerate-your-career-in-2025/"&gt;How to use AI to accelerate your career in 2025 (2025-09-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/1/how-data-scientists-can-master-life-sciences-and-software-skills-for-biotech-using-ultralearning/"&gt;How data scientists can master life sciences and software skills for biotech using ultralearning (2025-10-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/17/the-selfish-reason-to-do-your-best-work/"&gt;The selfish reason to do your best work (2025-12-17)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="data-science-practice-leadership"&gt;Data Science Practice &amp;amp; Leadership&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/10/a-practical-guide-to-securing-secrets-in-data-science-projects/"&gt;A practical guide to securing secrets in data science projects (2025-01-10)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/19/why-data-from-preclinical-biotech-lab-experiments-make-machine-learning-challenging/"&gt;Why data from preclinical biotech lab experiments make machine learning challenging (2025-01-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/31/pydata-bostoncambridge-talk-moderna-what-makes-an-agent/"&gt;PyData Boston/Cambridge Talk @ Moderna: What makes an agent? (2025-01-31)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/23/reliable-biological-data-requires-physical-quantities-not-statistical-artifacts/"&gt;Reliable biological data requires physical quantities, not statistical artifacts (2025-02-23)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/6/a-blueprint-for-data-driven-molecule-engineering/"&gt;A blueprint for data-driven molecule engineering (2025-03-06)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/16/the-art-of-finesse-as-a-data-scientist/"&gt;The art of finesse as a data scientist (2025-03-16)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/17/why-you-should-take-part-in-the-scipy-sprints/"&gt;Why you should take part in the SciPy sprints! (2025-03-17)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/2/how-to-standardize-data-science-ways-of-working-to-unlock-your-teams-creativity/"&gt;How to standardize Data Science ways of working to unlock your team's creativity (2025-04-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/3/bayesian-superiority-estimation-with-r2d2-priors-a-practical-guide-for-protein-screening/"&gt;Bayesian Superiority Estimation with R2D2 Priors: A Practical Guide for Protein Screening (2025-04-03)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/5/from-data-chaos-to-statistical-clarity-a-laboratory-transformation-story/"&gt;From data chaos to statistical clarity: A laboratory transformation story (2025-04-05)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/19/good-practices-for-ai-assisted-development-from-a-live-protein-calculator-demo/"&gt;Good practices for AI-assisted development from a live protein calculator demo (2025-04-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/8/why-im-excited-for-scipy-2025/"&gt;Why I'm excited for SciPy 2025! (2025-05-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/7/principles-for-using-ai-autodidactically/"&gt;Principles for using AI autodidactically (2025-06-07)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/27/build-your-own-tools/"&gt;Build your own tools! (2025-06-27)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/7/the-job-your-docs-need-to-do/"&gt;The job your docs need to do (2025-07-07)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/13/earn-the-privilege-to-use-automation/"&gt;Earn the privilege to use automation (2025-07-13)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/14/reflections-on-the-scipy-2025-conference/"&gt;Reflections on the SciPy 2025 Conference (2025-07-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/21/from-nerd-sniped-to-shipped-using-ai-as-a-thinking-tool/"&gt;From nerd-sniped to shipped using AI as a thinking tool (2025-07-21)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/6/stop-guessing-at-priors-r2d2s-automated-approach-to-bayesian-modeling/"&gt;Stop guessing at priors: R2D2's automated approach to Bayesian modeling (2025-08-06)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/15/data-scientists-arent-becoming-obsolete-in-the-llm-era/"&gt;Data scientists aren't becoming obsolete in the LLM era (2025-08-15)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/24/how-to-communicate-with-lab-scientists-when-youre-the-data-person/"&gt;How to communicate with lab scientists (when you're the data person) (2025-08-24)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/9/2/the-data-science-bootstrap-notes-a-major-upgrade-for-2025/"&gt;The Data Science Bootstrap Notes: A major upgrade for 2025 (2025-09-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/1/how-data-scientists-can-master-life-sciences-and-software-skills-for-biotech-using-ultralearning/"&gt;How data scientists can master life sciences and software skills for biotech using ultralearning (2025-10-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/19/how-to-expose-any-documentation-to-any-llm-agent/"&gt;How to expose any documentation to any LLM agent (2025-10-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/2/what-does-it-take-to-build-a-statistics-agent/"&gt;What does it take to build a statistics agent? (2025-12-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/17/the-selfish-reason-to-do-your-best-work/"&gt;The selfish reason to do your best work (2025-12-17)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="data-science-tooling"&gt;Data Science Tooling&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/4/what-makes-an-agent/"&gt;What makes an agent? (2025-01-04)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/10/a-practical-guide-to-securing-secrets-in-data-science-projects/"&gt;A practical guide to securing secrets in data science projects (2025-01-10)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/13/writing-at-the-speed-of-thought/"&gt;Writing at the speed of thought (2025-01-13)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/19/why-data-from-preclinical-biotech-lab-experiments-make-machine-learning-challenging/"&gt;Why data from preclinical biotech lab experiments make machine learning challenging (2025-01-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/31/pydata-bostoncambridge-talk-moderna-what-makes-an-agent/"&gt;PyData Boston/Cambridge Talk @ Moderna: What makes an agent? (2025-01-31)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/7/lightening-the-llamabot/"&gt;Lightening the LlamaBot (2025-02-07)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/17/let-me-ship-you-the-python-you-need/"&gt;Let me ship you the Python you need (2025-02-17)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/23/reliable-biological-data-requires-physical-quantities-not-statistical-artifacts/"&gt;Reliable biological data requires physical quantities, not statistical artifacts (2025-02-23)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/1/how-to-fix-pypi-upload-errors-related-to-license-metadata/"&gt;How to fix PyPI upload errors related to license metadata (2025-03-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/6/a-blueprint-for-data-driven-molecule-engineering/"&gt;A blueprint for data-driven molecule engineering (2025-03-06)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/17/why-you-should-take-part-in-the-scipy-sprints/"&gt;Why you should take part in the SciPy sprints! (2025-03-17)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/2/how-to-standardize-data-science-ways-of-working-to-unlock-your-teams-creativity/"&gt;How to standardize Data Science ways of working to unlock your team's creativity (2025-04-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/3/bayesian-superiority-estimation-with-r2d2-priors-a-practical-guide-for-protein-screening/"&gt;Bayesian Superiority Estimation with R2D2 Priors: A Practical Guide for Protein Screening (2025-04-03)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/5/from-data-chaos-to-statistical-clarity-a-laboratory-transformation-story/"&gt;From data chaos to statistical clarity: A laboratory transformation story (2025-04-05)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/8/wow-marimo/"&gt;Wow, Marimo! (2025-04-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/19/good-practices-for-ai-assisted-development-from-a-live-protein-calculator-demo/"&gt;Good practices for AI-assisted development from a live protein calculator demo (2025-04-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/26/wow-modal/"&gt;Wow, Modal! (2025-04-26)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/8/why-im-excited-for-scipy-2025/"&gt;Why I'm excited for SciPy 2025! (2025-05-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/24/supercharge-your-coding-agents-with-vscode-workspaces/"&gt;Supercharge your coding agents with VSCode workspaces (2025-05-24)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/25/the-invisible-polish-of-automatic-model-routing/"&gt;The invisible polish of automatic model routing (2025-05-25)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/14/rethinking-llm-interfaces-from-chatbots-to-contextual-applications/"&gt;Rethinking LLM interfaces, from chatbots to contextual applications (2025-06-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/27/build-your-own-tools/"&gt;Build your own tools! (2025-06-27)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/1/one-hour-and-eight-minutes-building-a-receipt-scanner-with-the-weirdest-tech-stack-imaginable/"&gt;One hour and eight minutes: Building a receipt scanner with the weirdest tech stack imaginable (2025-07-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/13/earn-the-privilege-to-use-automation/"&gt;Earn the privilege to use automation (2025-07-13)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/14/reflections-on-the-scipy-2025-conference/"&gt;Reflections on the SciPy 2025 Conference (2025-07-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/15/how-to-use-xarray-for-unified-laboratory-data-storage/"&gt;How to use xarray for unified laboratory data storage (2025-07-15)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/21/from-nerd-sniped-to-shipped-using-ai-as-a-thinking-tool/"&gt;From nerd-sniped to shipped using AI as a thinking tool (2025-07-21)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/6/stop-guessing-at-priors-r2d2s-automated-approach-to-bayesian-modeling/"&gt;Stop guessing at priors: R2D2's automated approach to Bayesian modeling (2025-08-06)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/15/data-scientists-arent-becoming-obsolete-in-the-llm-era/"&gt;Data scientists aren't becoming obsolete in the LLM era (2025-08-15)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/23/wicked-python-trickery-dynamically-patch-a-python-functions-source-code-at-runtime/"&gt;Wicked Python trickery - dynamically patch a Python function's source code at runtime (2025-08-23)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/9/2/the-data-science-bootstrap-notes-a-major-upgrade-for-2025/"&gt;The Data Science Bootstrap Notes: A major upgrade for 2025 (2025-09-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/4/how-to-teach-your-coding-agent-with-agentsmd/"&gt;How to teach your coding agent with AGENTS.md (2025-10-04)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/10/how-to-use-multiple-github-accounts-on-the-same-computer/"&gt;How to use multiple GitHub accounts on the same computer (2025-10-10)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/14/how-to-use-coding-agents-effectively/"&gt;How to Use Coding Agents Effectively (2025-10-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/18/a-practical-comparison-of-dspy-and-llamabot-for-structured-llm-applications/"&gt;A practical comparison of DSPy and LlamaBot for structured LLM applications (2025-10-18)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/19/how-to-expose-any-documentation-to-any-llm-agent/"&gt;How to expose any documentation to any LLM agent (2025-10-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/20/exploring-skills-vs-mcp-servers/"&gt;Exploring Skills vs MCP Servers (2025-10-20)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/28/use-coding-agents-to-write-marimo-notebooks/"&gt;Use coding agents to write Marimo notebooks (2025-10-28)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/8/safe-ways-to-let-your-coding-agent-work-autonomously/"&gt;Safe ways to let your coding agent work autonomously (2025-11-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/"&gt;How I Replaced 307 Lines of Agent Code with 4 Lines (2025-11-16)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/17/how-to-reference-code-across-repositories-with-coding-agents/"&gt;How to Reference Code Across Repositories with Coding Agents (2025-11-17)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/2/what-does-it-take-to-build-a-statistics-agent/"&gt;What does it take to build a statistics agent? (2025-12-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/10/productive-patterns-for-agent-assisted-programming/"&gt;Productive Patterns for Agent-Assisted Programming (2025-12-10)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="llms"&gt;LLMs&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/4/what-makes-an-agent/"&gt;What makes an agent? (2025-01-04)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/31/pydata-bostoncambridge-talk-moderna-what-makes-an-agent/"&gt;PyData Boston/Cambridge Talk @ Moderna: What makes an agent? (2025-01-31)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/7/lightening-the-llamabot/"&gt;Lightening the LlamaBot (2025-02-07)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/8/wow-marimo/"&gt;Wow, Marimo! (2025-04-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/26/wow-modal/"&gt;Wow, Modal! (2025-04-26)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/8/why-im-excited-for-scipy-2025/"&gt;Why I'm excited for SciPy 2025! (2025-05-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/24/supercharge-your-coding-agents-with-vscode-workspaces/"&gt;Supercharge your coding agents with VSCode workspaces (2025-05-24)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/25/the-invisible-polish-of-automatic-model-routing/"&gt;The invisible polish of automatic model routing (2025-05-25)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/7/principles-for-using-ai-autodidactically/"&gt;Principles for using AI autodidactically (2025-06-07)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/14/rethinking-llm-interfaces-from-chatbots-to-contextual-applications/"&gt;Rethinking LLM interfaces, from chatbots to contextual applications (2025-06-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/27/build-your-own-tools/"&gt;Build your own tools! (2025-06-27)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/1/one-hour-and-eight-minutes-building-a-receipt-scanner-with-the-weirdest-tech-stack-imaginable/"&gt;One hour and eight minutes: Building a receipt scanner with the weirdest tech stack imaginable (2025-07-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/13/earn-the-privilege-to-use-automation/"&gt;Earn the privilege to use automation (2025-07-13)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/14/reflections-on-the-scipy-2025-conference/"&gt;Reflections on the SciPy 2025 Conference (2025-07-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/21/from-nerd-sniped-to-shipped-using-ai-as-a-thinking-tool/"&gt;From nerd-sniped to shipped using AI as a thinking tool (2025-07-21)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/15/data-scientists-arent-becoming-obsolete-in-the-llm-era/"&gt;Data scientists aren't becoming obsolete in the LLM era (2025-08-15)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/23/wicked-python-trickery-dynamically-patch-a-python-functions-source-code-at-runtime/"&gt;Wicked Python trickery - dynamically patch a Python function's source code at runtime (2025-08-23)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/9/1/how-to-use-ai-to-accelerate-your-career-in-2025/"&gt;How to use AI to accelerate your career in 2025 (2025-09-01)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/4/how-to-teach-your-coding-agent-with-agentsmd/"&gt;How to teach your coding agent with AGENTS.md (2025-10-04)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/14/how-to-use-coding-agents-effectively/"&gt;How to Use Coding Agents Effectively (2025-10-14)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/18/a-practical-comparison-of-dspy-and-llamabot-for-structured-llm-applications/"&gt;A practical comparison of DSPy and LlamaBot for structured LLM applications (2025-10-18)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/19/how-to-expose-any-documentation-to-any-llm-agent/"&gt;How to expose any documentation to any LLM agent (2025-10-19)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/20/exploring-skills-vs-mcp-servers/"&gt;Exploring Skills vs MCP Servers (2025-10-20)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/28/use-coding-agents-to-write-marimo-notebooks/"&gt;Use coding agents to write Marimo notebooks (2025-10-28)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/8/safe-ways-to-let-your-coding-agent-work-autonomously/"&gt;Safe ways to let your coding agent work autonomously (2025-11-08)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/"&gt;How I Replaced 307 Lines of Agent Code with 4 Lines (2025-11-16)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/17/how-to-reference-code-across-repositories-with-coding-agents/"&gt;How to Reference Code Across Repositories with Coding Agents (2025-11-17)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/2/what-does-it-take-to-build-a-statistics-agent/"&gt;What does it take to build a statistics agent? (2025-12-02)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/10/productive-patterns-for-agent-assisted-programming/"&gt;Productive Patterns for Agent-Assisted Programming (2025-12-10)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="all-blog-posts"&gt;All blog posts&lt;/h2&gt;&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;th&gt;Categories&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2025-01-04&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/4/what-makes-an-agent/"&gt;What makes an agent?&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-01-10&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/10/a-practical-guide-to-securing-secrets-in-data-science-projects/"&gt;A practical guide to securing secrets in data science projects&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-01-13&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/13/writing-at-the-speed-of-thought/"&gt;Writing at the speed of thought&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-01-19&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/19/why-data-from-preclinical-biotech-lab-experiments-make-machine-learning-challenging/"&gt;Why data from preclinical biotech lab experiments make machine learning challenging&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-01-31&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/1/31/pydata-bostoncambridge-talk-moderna-what-makes-an-agent/"&gt;PyData Boston/Cambridge Talk @ Moderna: What makes an agent?&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling, Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-02-07&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/7/lightening-the-llamabot/"&gt;Lightening the LlamaBot&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-02-17&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/17/let-me-ship-you-the-python-you-need/"&gt;Let me ship you the Python you need&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-02-23&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/2/23/reliable-biological-data-requires-physical-quantities-not-statistical-artifacts/"&gt;Reliable biological data requires physical quantities, not statistical artifacts&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Biology &amp;amp; Chemistry, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-03-01&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/1/how-to-fix-pypi-upload-errors-related-to-license-metadata/"&gt;How to fix PyPI upload errors related to license metadata&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-03-06&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/6/a-blueprint-for-data-driven-molecule-engineering/"&gt;A blueprint for data-driven molecule engineering&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-03-16&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/16/the-art-of-finesse-as-a-data-scientist/"&gt;The art of finesse as a data scientist&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-03-17&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/3/17/why-you-should-take-part-in-the-scipy-sprints/"&gt;Why you should take part in the SciPy sprints!&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-04-02&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/2/how-to-standardize-data-science-ways-of-working-to-unlock-your-teams-creativity/"&gt;How to standardize Data Science ways of working to unlock your team's creativity&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-04-03&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/3/bayesian-superiority-estimation-with-r2d2-priors-a-practical-guide-for-protein-screening/"&gt;Bayesian Superiority Estimation with R2D2 Priors: A Practical Guide for Protein Screening&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-04-05&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/5/from-data-chaos-to-statistical-clarity-a-laboratory-transformation-story/"&gt;From data chaos to statistical clarity: A laboratory transformation story&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-04-08&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/8/wow-marimo/"&gt;Wow, Marimo!&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling, LLMs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-04-19&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/19/good-practices-for-ai-assisted-development-from-a-live-protein-calculator-demo/"&gt;Good practices for AI-assisted development from a live protein calculator demo&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-04-26&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/4/26/wow-modal/"&gt;Wow, Modal!&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling, LLMs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-05-08&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/8/why-im-excited-for-scipy-2025/"&gt;Why I'm excited for SciPy 2025!&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership, Data Science Tooling, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-05-24&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/24/supercharge-your-coding-agents-with-vscode-workspaces/"&gt;Supercharge your coding agents with VSCode workspaces&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-05-25&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/5/25/the-invisible-polish-of-automatic-model-routing/"&gt;The invisible polish of automatic model routing&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-06-07&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/7/principles-for-using-ai-autodidactically/"&gt;Principles for using AI autodidactically&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-06-14&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/14/rethinking-llm-interfaces-from-chatbots-to-contextual-applications/"&gt;Rethinking LLM interfaces, from chatbots to contextual applications&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-06-27&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/6/27/build-your-own-tools/"&gt;Build your own tools!&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-07-01&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/1/one-hour-and-eight-minutes-building-a-receipt-scanner-with-the-weirdest-tech-stack-imaginable/"&gt;One hour and eight minutes: Building a receipt scanner with the weirdest tech stack imaginable&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-07-07&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/7/the-job-your-docs-need-to-do/"&gt;The job your docs need to do&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-07-13&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/13/earn-the-privilege-to-use-automation/"&gt;Earn the privilege to use automation&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-07-14&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/14/reflections-on-the-scipy-2025-conference/"&gt;Reflections on the SciPy 2025 Conference&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership, Data Science Tooling, Biology &amp;amp; Chemistry, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-07-15&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/15/how-to-use-xarray-for-unified-laboratory-data-storage/"&gt;How to use xarray for unified laboratory data storage&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-07-21&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/7/21/from-nerd-sniped-to-shipped-using-ai-as-a-thinking-tool/"&gt;From nerd-sniped to shipped using AI as a thinking tool&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-08-06&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/6/stop-guessing-at-priors-r2d2s-automated-approach-to-bayesian-modeling/"&gt;Stop guessing at priors: R2D2's automated approach to Bayesian modeling&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-08-15&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/15/data-scientists-arent-becoming-obsolete-in-the-llm-era/"&gt;Data scientists aren't becoming obsolete in the LLM era&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Practice &amp;amp; Leadership, Data Science Tooling, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-08-23&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/23/wicked-python-trickery-dynamically-patch-a-python-functions-source-code-at-runtime/"&gt;Wicked Python trickery - dynamically patch a Python function's source code at runtime&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-08-24&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/8/24/how-to-communicate-with-lab-scientists-when-youre-the-data-person/"&gt;How to communicate with lab scientists (when you're the data person)&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Biology &amp;amp; Chemistry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-09-01&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/9/1/how-to-use-ai-to-accelerate-your-career-in-2025/"&gt;How to use AI to accelerate your career in 2025&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-09-02&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/9/2/the-data-science-bootstrap-notes-a-major-upgrade-for-2025/"&gt;The Data Science Bootstrap Notes: A major upgrade for 2025&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-01&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/1/how-data-scientists-can-master-life-sciences-and-software-skills-for-biotech-using-ultralearning/"&gt;How data scientists can master life sciences and software skills for biotech using ultralearning&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Practice &amp;amp; Leadership, Biology &amp;amp; Chemistry, Career Advice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-04&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/4/how-to-teach-your-coding-agent-with-agentsmd/"&gt;How to teach your coding agent with AGENTS.md&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-10&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/10/how-to-use-multiple-github-accounts-on-the-same-computer/"&gt;How to use multiple GitHub accounts on the same computer&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-14&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/14/how-to-use-coding-agents-effectively/"&gt;How to Use Coding Agents Effectively&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-18&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/18/a-practical-comparison-of-dspy-and-llamabot-for-structured-llm-applications/"&gt;A practical comparison of DSPy and LlamaBot for structured LLM applications&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-19&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/19/how-to-expose-any-documentation-to-any-llm-agent/"&gt;How to expose any documentation to any LLM agent&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling, Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-20&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/20/exploring-skills-vs-mcp-servers/"&gt;Exploring Skills vs MCP Servers&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-10-28&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/10/28/use-coding-agents-to-write-marimo-notebooks/"&gt;Use coding agents to write Marimo notebooks&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-11-08&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/8/safe-ways-to-let-your-coding-agent-work-autonomously/"&gt;Safe ways to let your coding agent work autonomously&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-11-16&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/"&gt;How I Replaced 307 Lines of Agent Code with 4 Lines&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-11-17&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/11/17/how-to-reference-code-across-repositories-with-coding-agents/"&gt;How to Reference Code Across Repositories with Coding Agents&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Data Science Tooling, LLMs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-12-02&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/2/what-does-it-take-to-build-a-statistics-agent/"&gt;What does it take to build a statistics agent?&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling, Biology &amp;amp; Chemistry, Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-12-10&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/10/productive-patterns-for-agent-assisted-programming/"&gt;Productive Patterns for Agent-Assisted Programming&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LLMs, Data Science Tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-12-17&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ericmjl.github.io/blog/2025/12/17/the-selfish-reason-to-do-your-best-work/"&gt;The selfish reason to do your best work&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Career Advice, Data Science Practice &amp;amp; Leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
</content></entry><entry><title>The selfish reason to do your best work</title><link href="https://ericmjl.github.io/blog/2025/12/17/the-selfish-reason-to-do-your-best-work/" rel="alternate"/><updated>2025-12-17T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:0e85eb9c-a072-3e9f-8de1-dbcb432865e4</id><content type="html">&lt;p&gt;I’ve been thinking a lot about career lately. This has been a pretty lean year for biotech; we've seen ups and downs at Moderna and across the industry. So, I want to offer a word of encouragement and a philosophy on work that I hope can be useful for you, regardless of where you are in your journey.&lt;/p&gt;
&lt;p&gt;It starts with a reframing of &lt;em&gt;why&lt;/em&gt; we work.&lt;/p&gt;
&lt;h3 id="do-your-best-work-for-yourself"&gt;Do your best work for yourself&lt;/h3&gt;&lt;p&gt;I know there is a lot of sentiment going around the internet right now about "acting your wage"—limiting your effort to exactly what you are paid for—or doing the bare minimum. The logic goes: &lt;em&gt;Why should I care about doing my best for my job if my company doesn't care for me?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I get where that sentiment comes from. But I want to redirect your attention a little bit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;You do not have to do your best work for your company. You should do your best work for &lt;em&gt;yourself&lt;/em&gt; at the company you’re at.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It just so happens that the company will benefit, but you should treat your effort as an investment in your own professional instincts and habits. Sooner or later, you may not be at that company. You actually don't really have to care about the entity itself. If you don't care about the people who work at your workplace, no one is compelling you to. (Though, if you do happen to like your colleagues—which is true for me where I work—then that’s all good.)&lt;/p&gt;
&lt;p&gt;But even if you can’t find much to be inspired by, do your best work anyway. Why? Because you are building the person you will be in five or ten years.&lt;/p&gt;
&lt;p&gt;President Obama once gave this advice to young interns: "Don't ask for the plum assignments. Just knock out everything you're doing." I guarantee you someone will notice. Even if no one at your current company notices, if you build a track record of quality, people &lt;em&gt;outside&lt;/em&gt; will notice.&lt;/p&gt;
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/YNY4UFaHbP4?si=_tkyjLX25IWCfF0L" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen&gt;&lt;/iframe&gt;&lt;p&gt;Your reputation precedes you. It is the one thing you accumulate over time that serves as a form of wealth that can never be taken away from you. Only you can lose it. To borrow a phrase from Jocko Willink, this is "Extreme Ownership"—taking total responsibility for your world. Yes, circumstances happen, but if you guard your reputation well, it is yours to keep.&lt;/p&gt;
&lt;p&gt;Think about it: Who knows where you will be ten years down the road? If you are a software engineer or data scientist now, in five years you might be a Director. You’re going to be calling the shots. If you don't take the time &lt;em&gt;now&lt;/em&gt; to practice making decisions, witnessing judgment calls, and battle-testing your engineering foresight, will you be ready?&lt;/p&gt;
&lt;p&gt;I had a former teammate who worked under me at Moderna, &lt;a href="https://www.linkedin.com/in/arkadij-kummer-78b249b9"&gt;Arkadij Kummer&lt;/a&gt;. He’s now the CTO of a startup—a title I haven't even held. I saw him put in the effort to develop the strategic thinking patterns that helped him get the skills he needed to lead a tech organization. He sought out opportunities to practice making judgment calls and owning the outcomes. You have to get those reps in early, or you won’t be ready for what happens later.&lt;/p&gt;
&lt;p&gt;So, if you are working at a company you want to leave: do not give up on investing in yourself. Fortune favors the prepared.&lt;/p&gt;
&lt;h3 id="resilience-is-also-an-investment"&gt;Resilience is also an investment&lt;/h3&gt;&lt;p&gt;Building this "career wealth" isn't just about technical execution. It's also about how you handle failure. And here is the second word of encouragement I want to offer: &lt;strong&gt;Everyone will make a blunder at some point.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Part of growing up—and part of investing in your own character—is owning up to those mistakes, being proactive about remedying them, and graciously accepting help.&lt;/p&gt;
&lt;p&gt;I have been there. I once wrote a very scathing internal blog post about my leadership at a previous company. Looking back, I sometimes think of it as the darkest weekend of my career.&lt;/p&gt;
&lt;p&gt;I was a guy who had just finished school, entered the workforce, and within three years decided I knew better than people who had done their jobs for twenty-odd years. I'm not saying it's impossible that I was right, but it was pretty presumptuous. I lashed out at other groups that I thought were incompetent, effectively attacking my own team.&lt;/p&gt;
&lt;p&gt;The leadership team responded with extreme grace. They knew the problems I outlined were real, but they also saw a junior person who hadn't picked his battles wisely.&lt;/p&gt;
&lt;p&gt;We have limited energy and limited ability to focus. There are only so many battles we can handle simultaneously. I chose a bad battle to fight.&lt;/p&gt;
&lt;p&gt;That experience prompted deep reflection. I decided to lean back into that first philosophy: I was going to go back and do a good job. Not necessarily because I was feeling proud of the company at that moment, but because I recognized that if I ever wanted to be the type of person with the authority to change things, I needed to be the best version of myself first. I needed to learn how to handle authority and how to elevate the people around me.&lt;/p&gt;
&lt;p&gt;If you are in a situation where you’ve made a mistake, the best thing you can do for your reputation is to own it. Propose an action to rectify it. Move on.&lt;/p&gt;
&lt;p&gt;Most mistakes are not unforgivable. There is a classic business story about Tom Watson, the founder of IBM, involving a subordinate who made a mistake that cost the company \$600,000. Whether the amount was really \$600,000 or not, the story has a lesson that rings true. The man walked into Watson's office expecting to be fired. Watson reportedly replied, "No, I just spent \$600,000 training you. Why would I want to fire you?"&lt;/p&gt;
&lt;p&gt;Most people will understand. Yes, there will be delays. That is just the world we inhabit. Own the mistake, improve the process, and keep making it better. Again, not primarily for the company, but because you are preparing yourself to lead with integrity and compassion.&lt;/p&gt;
&lt;h3 id="the-wealth-that-remains"&gt;The wealth that remains&lt;/h3&gt;&lt;p&gt;I don't want you to become the whiner or the complainer. That is a habit that will stick with you for the rest of your life.&lt;/p&gt;
&lt;p&gt;Instead, I want you to play the long game. Don't let temporary frustrations dictate your long-term growth. Anything worth doing is going to be difficult. If it was easy, the reward would be fleeting. Even for the ultra-rich, like Sergey Brin, eventually the yacht gets boring. He returned to active coding at Google to work on AI because, fundamentally, humans are wired to find satisfaction in building something meaningful.&lt;/p&gt;
&lt;p&gt;So, aspire to greatness. Not just to gain a title or a promotion, but because in the process, you accumulate a wealth that cannot be taken away from you: You will have the satisfaction of mastery. You will have a battle-tested character. You will have a reputation that opens doors before you even knock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Do your best work.&lt;/strong&gt; It’s the best investment you’ll ever make. And it'll be the most selfless gift you give to yourself and others.&lt;/p&gt;
</content></entry><entry><title>Productive Patterns for Agent-Assisted Programming</title><link href="https://ericmjl.github.io/blog/2025/12/10/productive-patterns-for-agent-assisted-programming/" rel="alternate"/><updated>2025-12-10T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:fe2a71e9-46d4-3221-a3fa-3f7dd32c2786</id><content type="html">&lt;p&gt;I've been using coding agents for a while now, and I've learned a few patterns that make the experience much more productive. The thing is, a lot of these "productive patterns" aren't being shared enough—they're more like folk knowledge that you can only really pick up by watching someone else do their work live. I decided to write this blog post to kickstart conversations about the matter. Here's what works for me.&lt;/p&gt;
&lt;h2 id="build-a-detailed-plan-with-ai"&gt;Build a detailed plan with AI&lt;/h2&gt;&lt;p&gt;Before jumping into implementation, spend time building a detailed plan with your AI assistant. Iterate 2-3 times over the plan, checking every detail. You want the ability to see in your head what the code might look like—just a "fat finger sketch" of the implementation.&lt;/p&gt;
&lt;p&gt;The plan should include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Details on implementation&lt;/li&gt;
&lt;li&gt;How to test (this is the most important part)&lt;/li&gt;
&lt;li&gt;Documentation plan&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="do-docs-and-tests-first"&gt;Do docs and tests first&lt;/h2&gt;&lt;p&gt;Humans usually adhere to test-driven development if you're a software engineer, or exploration-driven software builds if you're more of a data scientist. Because of the sequential nature of generative AI, it's advantageous to instruct AI to do the docs and tests first before the implementation. This is a complex conditional probability problem. If the tests and docs are written first, the implementation has to satisfy those constraints, which leads to better code.&lt;/p&gt;
&lt;p&gt;The test plan should include instructions on how to run tests using command line tools. Don't assume the AI knows your project's specific testing setup.&lt;/p&gt;
&lt;h2 id="use-agents-md-as-your-repo-s-ai-university"&gt;Use AGENTS.md as your repo's AI university&lt;/h2&gt;&lt;p&gt;AGENTS.md is a great place to store the specific instructions that you need for the repo. For example, AI will tend to write &lt;code&gt;python -m ...&lt;/code&gt; as a shell command, but if I'm running a pixi project, it's better to always run &lt;code&gt;pixi run python ...&lt;/code&gt; instead. Treat AGENTS.md as the AI's university of your particular repo; it's where you encode all the project-specific knowledge that the AI needs to work effectively.&lt;/p&gt;
&lt;h2 id="control-the-pace-of-the-agent"&gt;Control the pace of the agent&lt;/h2&gt;&lt;p&gt;Know the default behavior of your agent; it may be over-eager to do lots of things. You can pace the coding agent by asking it to "slow down, walk me through the changes one at a time, starting with the most important ones first." This helps you maintain control and review changes as they happen, rather than being overwhelmed by a massive diff.&lt;/p&gt;
&lt;h2 id="leverage-local-and-command-line-tools"&gt;Leverage local and command line tools&lt;/h2&gt;&lt;p&gt;You can use local and command line tools to your advantage! Here are some examples:&lt;/p&gt;
&lt;p&gt;Firstly, the GitHub CLI (&lt;code&gt;gh&lt;/code&gt;) can be used to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Store plans on GitHub as issues first (a matter of taste—you can avoid cluttering up your local filesystem)&lt;/li&gt;
&lt;li&gt;Pull GitHub Actions logs&lt;/li&gt;
&lt;li&gt;Call out to the GitHub API for other general tasks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Environment management:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;pixi run&lt;/code&gt; ensures you're always running within the correct Python environment&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uvx marimo check&lt;/code&gt; lets me check that marimo notebooks are syntactically valid&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uv run notebook.py&lt;/code&gt; lets me run notebooks as scripts to check outputs&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uvx marimo export&lt;/code&gt; lets me export marimo notebooks as markdown&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Linting and quality:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;markdownlint&lt;/code&gt; runs on every edit of markdown files so you never have markdown linting issues&lt;/li&gt;
&lt;li&gt;Get AI to "commit relevant files and fix issues raised by pre-commit hooks"&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Let agents use CLI tools and read outputs directly so that you don't have to switch between windows copying and pasting things manually.&lt;/p&gt;
&lt;h2 id="let-agents-write-temporary-tools"&gt;Let agents write temporary tools&lt;/h2&gt;&lt;p&gt;Coding agents can write their own temporary tools inside &lt;code&gt;.py&lt;/code&gt; files. Encourage coding agents to do that to test that what it wrote works on-the-fly. This is a great way to validate code before integrating it into your main codebase.&lt;/p&gt;
&lt;p&gt;You can even experiment with self-improving agents: if it detects you correcting its action, it should auto-update AGENTS.md with what is the correct thing to do. I haven't fully battle-tested this yet, but you can write an "AI constitution" at the top of AGENTS.md that instructs the agent to learn from corrections by &lt;em&gt;remembering them inside AGENTS.md&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id="develop-your-own-tools"&gt;Develop your own tools&lt;/h2&gt;&lt;p&gt;Isabel Zimmerman mentioned this in her keynote talk: develop your own tools. Here are some examples of my own:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Personal MCP productivity server&lt;/strong&gt;: gives me prompts that I can take from project to project, so I don't have to keep copying/pasting them&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shell aliases&lt;/strong&gt;: &lt;code&gt;gacp&lt;/code&gt; lets me run &lt;code&gt;git add . &amp;amp;&amp;amp; git commit &amp;amp;&amp;amp; git push&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LlamaBot git hooks&lt;/strong&gt;: auto-writes commit messages for me&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These custom tools compound over time and make your workflow significantly more efficient.&lt;/p&gt;
&lt;p&gt;These patterns have made my agent-assisted programming much more productive. Treat the AI as a collaborator that needs clear instructions, proper context, and the right tools to work effectively. Start with a good plan, control the pace, and build tools that make the whole process smoother.&lt;/p&gt;
&lt;p&gt;What patterns have you discovered? I'd love to hear what works for you—let's make this folk knowledge more accessible to everyone.&lt;/p&gt;
</content></entry><entry><title>What does it take to build a statistics agent?</title><link href="https://ericmjl.github.io/blog/2025/12/2/what-does-it-take-to-build-a-statistics-agent/" rel="alternate"/><updated>2025-12-02T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:3edf8f61-a0a7-3f5c-b699-ee70ca69638e</id><content type="html">&lt;p&gt;Within research organizations at most pharma and biotech companies, professionally-trained statisticians are often staffed at extremely low ratios relative to the number of lab scientists. By rough Fermi estimation, I'd hazard a guess that ratios anywhere from 1:10 to 1:100 are plausible, meaning most researchers have limited access to statistical expertise when they need it most, during experiment design. This statistician shortage creates a critical bottleneck in experimental design, power calculations, and biostatistical consultation—areas where proper statistical guidance can prevent costly mistakes and improve research reproducibility.&lt;/p&gt;
&lt;p&gt;This creates a costly problem. When statisticians aren't available, researchers fall back to what I call "folk statistics" - the kind you learn by immersion in a lab, or from 1-2 graduate lectures hidden within broader "laboratory methods" or "computational methods" classes. I know this because I practiced folk statistics myself in the life sciences, blindly following rules like "just do n=3" or "just use the t-test with your count data" without understanding the statistical reasoning behind these choices.&lt;/p&gt;
&lt;p&gt;The consequences are documented in stark numbers. Amgen scientists attempted to reproduce 53 landmark preclinical papers and failed in 47 cases (89%)—even after contacting original authors and exchanging reagents. Bayer's internal validation found only 20-25% of studies "completely in line" with original publications. These studies consistently identified poor experimental design and inadequate statistical analysis as major contributors. &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC4461318/"&gt;Freedman et al. (2015)&lt;/a&gt; estimated $28 billion annually spent on irreproducible preclinical research in the United States alone.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://cdn.ncbi.nlm.nih.gov/pmc/blobs/ccea/4461318/0e3c7aea2b45/pbio.1002165.g002.jpg" alt="Breakdown of causes of preclinical irreproducibility from Freedman et al. (2015). Study design accounts for 27.6% of irreproducibility."&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Breakdown of causes of preclinical irreproducibility from Freedman et al. (2015). Study design accounts for 27.6% of irreproducibility.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;At the individual experiment level, this translates to teams &lt;strong&gt;throwing out hard-won experimental data&lt;/strong&gt; that can cost anywhere from thousands to hundreds of thousands of dollars to collect per round, &lt;strong&gt;wasting up to millions of dollars downstream&lt;/strong&gt; by basing decisions on poorly-collected data, and &lt;strong&gt;missing opportunities&lt;/strong&gt; to set up machine learners with high quality laboratory data that could shortcut the amount of laboratory experimentation needed.&lt;/p&gt;
&lt;p&gt;I took one semester of graduate-level biostatistics, then a decade of self-study in Bayesian statistics, followed by professional work where accurate estimation was critical—whether estimating half-life of a molecule, binding affinity of an antibody, or other performance properties. Through this journey, I no longer trust folk statistics. Folk statistics relies on faulty assumptions—like "n=3 is all you'll really need," "use the t-test for count data," or "calculate the SEM and don't show the SD"—which influence bad decision-making when people don't know better. Once you see how these assumptions break down and lead to wrong conclusions, you can't unsee it. Quantities like half-life, binding affinity, and other performance properties need to be accurately estimated through proper experimental design and statistically-informed mechanistic modeling.&lt;/p&gt;
&lt;p&gt;Statisticians are expensive, but they're also 100% critical for generating high quality, high fidelity data. Their role at the experiment design phase is usually that of a consultant, asking probing questions to ensure experiments are designed with good controls, confounders are accounted for, and the right statistical models are chosen. The question is: &lt;strong&gt;can we scale this expertise?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Not to replace statisticians, but to level up the organizational statistical practice &lt;em&gt;before&lt;/em&gt; researchers check in with a professionally-trained stats person. If lab scientists can think through their experimental designs more rigorously beforehand - understanding power calculations, considering confounders, planning proper controls - then the conversations they have with statisticians can be elevated. Instead of starting from scratch, they can engage in more sophisticated discussions about design trade-offs, model selection, and advanced statistical considerations. In turn, this amplifies the value of the statistician's time and improves outcomes for everyone.&lt;/p&gt;
&lt;p&gt;I was inspired by Dr. Emi Tanaka's &lt;a href="https://emitanaka.org/slides/AASC2024/#/title-slide"&gt;slides on extracting elements of statistical experiment design using LLMs&lt;/a&gt;, which showed how we can extract structured information like response variables, treatments, experimental units, design types, replicate structure, and controls. I decided to take a stab at building something that could do more than just extract information—something that could actually consult on experiment design.&lt;/p&gt;
&lt;p&gt;And so &lt;code&gt;stats-agents&lt;/code&gt; was born: an AI-powered statistics agent for experiment design consultation. Here's how I designed and evaluated this domain-specific AI agent.&lt;/p&gt;
&lt;p&gt;As a preface, I initially explored the ReAct pattern but &lt;a href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/"&gt;switched to PocketFlow&lt;/a&gt;, a minimalist graph-based framework that replaced 307 lines of agent orchestration code with just 4 lines. This graph-based approach brought clarity, modularity, and made the execution flow explicit—exactly what I needed for building a robust statistics agent.&lt;/p&gt;
&lt;p&gt;So how did I go about building this agent?&lt;/p&gt;
&lt;p&gt;Deeply influenced by Clayton Christensen's books, I actually started with "what's the job for this agent to be done?" I initially considered building a single agent that could handle both experiment design consultation and statistical analysis of collected data. However, I quickly realized these are fundamentally different phases with different goals, tools, and interaction patterns.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;experiment design phase&lt;/strong&gt; is consultative and exploratory - it's about asking questions, understanding constraints, identifying potential issues, and helping researchers think through their design &lt;em&gt;before&lt;/em&gt; data collection. The &lt;strong&gt;analysis phase&lt;/strong&gt; is more technical - it's about taking collected data and building statistical models to estimate quantities of interest.&lt;/p&gt;
&lt;p&gt;I decided to focus the agent on the design phase only. This separation of concerns made the agent cleaner, less confusing, and allowed it to be optimized for its specific purpose: being an inquisitive, consultative partner during experiment design. The analysis phase would be handled separately (or by a different agent) with its own tools and prompting strategies.&lt;/p&gt;
&lt;p&gt;So I defined the agent's job description (JD) as: "an agent that will provide critique on experiment designs, suggest modifications, and help researchers think through their experimental design before data collection". It sounded oddly like a human's job description for a real job, except more specific. Notice, however, that the JD leaves room for a real human, in that no accountability for outcomes is placed on the agent, a human statistician still needs to review the work, just as we wrote above.&lt;/p&gt;
&lt;p&gt;With the job scope defined, I turned to designing the tools the agent would need.&lt;/p&gt;
&lt;p&gt;The first tool I gave was &lt;code&gt;critique_experiment_design&lt;/code&gt; - a tool that provides comprehensive critique of experiment designs, identifying potential flaws, biases, weaknesses, and areas for improvement. This tool considers multiple angles including biological, statistical, and practical constraints. The agent can use this to help researchers identify issues in their designs before they collect data.&lt;/p&gt;
&lt;p&gt;The second tool I gave was one I previously wrote about: the ability to execute code (&lt;code&gt;write_and_execute_code&lt;/code&gt;). I wanted this available to answer questions like "what should my data table look like?"&lt;/p&gt;
&lt;p&gt;I've noticed that having a sample data table in front of us when discussing experimental designs is incredibly clarifying—it cuts through abstract confusion toward concrete understanding. This tool enables the agent to generate sample data tables, perform power calculations, create plate map visualizations, and other dynamic analyses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Important security note&lt;/strong&gt;: The ability to execute arbitrary Python code is powerful but also dangerous. An agent that can execute code can delete files, modify system configurations, access sensitive data, make network requests, and more. For any production deployment, this agent must run in a containerized environment with strict isolation, resource limits, and no access to secrets or credentials. This isn't optional - it's a cybersecurity requirement! For my development and testing, I ran it in a controlled environment on my own machine, but production deployment would require proper containerization.&lt;/p&gt;
&lt;p&gt;With the tools defined, the next challenge was evaluation: how do you know if the agent is actually working? This turned out to be more complex than I initially expected.&lt;/p&gt;
&lt;p&gt;The evaluation process had two distinct phases: an exploration-guided MVP phase (vibes-driven) and a post-MVP systematic testing phase.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 1: MVP Development (Vibes-Driven)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;During the MVP phase, I defined a fixed conversation path and repeatedly tested it manually in a Marimo notebook's chat UI. I tested with a sequence like: asking for design critique, providing an experiment description, requesting power calculations, and asking for sample data tables. As I found errors, I fixed them immediately, contrary to evaluation best practices, but appropriate for this exploratory phase. The tool definitions weren't settled until this phase was complete.&lt;/p&gt;
&lt;p&gt;I also used Cursor (a coding agent) to help diagnose issues, explore multiple solutions, and get different perspectives before committing to fixes. This "multiple AI opinions before committing" pattern follows a similar philosophy to Geoffrey Litt's &lt;a href="https://www.geoffreylitt.com/2025/10/24/code-like-a-surgeon"&gt;"Code like a surgeon"&lt;/a&gt; approach: spike out an attempt at a big change, review it as a sketch of where to go, and often you won't use the result directly—but it helps you understand the problem space better.&lt;/p&gt;
&lt;p&gt;Rather than accepting the first AI suggestion, I'd ask multiple questions, explore several solution approaches, and understand the trade-offs before making an informed decision. When debugging complex issues like the closure vs. shared state problem, I'd often ask multiple models the same question to see if they'd converge on the same diagnosis—if different LLMs independently arrived at the same answer, that was a good sign the solution was on the right track. This led to better architecture decisions and fewer instances of "I wish I had done it differently."&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 2: Systematic Evaluation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Once the baseline behavior was satisfactory, I moved to systematic evaluation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Benchmark prompt&lt;/strong&gt;: I created a "perfect prompt" document (&lt;code&gt;experiment_design_for_power_calc.md&lt;/code&gt;) that served as my regression test suite. This complete, detailed experiment design specification should trigger specific agent behaviors, such as asking the right questions, performing power calculations correctly, providing contextual explanations. Every time I modified the system prompt or tools, I'd run this same benchmark and check: did it still work? Without this, I found myself making changes that broke things in subtle ways, or losing track of what "good" behavior even looked like. &lt;em&gt;The benchmark prompt became my north star, a concrete example of the agent working as intended.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Synthetic test generation&lt;/strong&gt;: Once I had a reliable benchmark, I expanded to systematic evaluation with variation. Starting with five examples from Tanaka's slide deck, I used the methodology from Shreya Shankar and Hamel Husain (researchers who developed systematic approaches for LLM evaluation) to generate synthetic chat examples by selecting 1-3 axes of variation. I chose experimental domain (biotech vs. agriculture) as the primary axis, while varying the statistical expertise level of the simulated user. This process generated dozens of conversation traces that, while not exhaustive, represented draws from my constrained prior belief about likely conversations.&lt;/p&gt;
&lt;p&gt;Systematic evaluation with these varied conversation traces revealed patterns I never would have noticed through manual testing. Through these traces, I identified three major categories of failure modes that needed to be addressed.&lt;/p&gt;
&lt;h2 id="failure-mode-1-multi-step-execution-breakdown"&gt;Failure mode 1: Multi-step execution breakdown&lt;/h2&gt;&lt;p&gt;The first major issue I encountered was that the agent couldn't chain tool calls effectively. When the agent executed code to perform a power calculation, it would store the result in a variable like &lt;code&gt;mtt_power_analysis_result&lt;/code&gt;. But when it tried to analyze that result in a subsequent tool call, it would fail with a &lt;code&gt;NameError&lt;/code&gt; - the variable simply wasn't accessible.&lt;/p&gt;
&lt;p&gt;The root cause was subtle: the code execution tool (&lt;code&gt;write_and_execute_code_wrapper&lt;/code&gt;) was using a closure variable that captured the notebook's globals at initialization time. However, the agent framework (AgentBot) stores results in a separate shared dictionary (&lt;code&gt;shared["globals_dict"]&lt;/code&gt;). These two dictionaries were disconnected; think of them as two separate notebooks that couldn't see each other's variables. So when the agent created a variable in one tool call, it wasn't visible to the next.&lt;/p&gt;
&lt;p&gt;The fix required connecting them: I modified the code execution tool to accept an optional &lt;code&gt;_globals_dict&lt;/code&gt; parameter. When provided, it uses the agent's shared dictionary instead of its own isolated one. This allows results from one tool call to be accessible in subsequent calls, enabling true multi-step workflows where the agent can build on previous results.&lt;/p&gt;
&lt;h2 id="failure-mode-2-display-formatting-and-contextual-output"&gt;Failure mode 2: Display formatting and contextual output&lt;/h2&gt;&lt;p&gt;The second category of issues involved both technical display problems and behavioral output quality. When the agent returned Python objects (DataFrames, matplotlib figures, etc.), they weren't displaying properly in the Marimo chat interface. The agent would return a dictionary with these objects, but Marimo's chat UI doesn't automatically render matplotlib &lt;code&gt;Figure&lt;/code&gt; objects embedded in dictionaries.&lt;/p&gt;
&lt;p&gt;But there was a deeper behavioral problem: the agent was dumping DataFrames and plots without any explanatory text. I'd ask for a power analysis, and the agent would return a raw DataFrame with numbers -- no context, no interpretation, no explanation of what I was looking at. My sense of taste rebelled. This wasn't just a technical problem, it was an aesthetic one. I wanted a polished, consultant-like experience where explanatory text naturally flows between objects, not a data dump.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sometimes the best technical solutions come from caring about how things look and feel.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I solved both problems through a combination of technical infrastructure and prompt engineering. With AI assistance, I created a formatter system that processes dictionaries containing text and objects, converting strings to markdown and passing objects through for native display. This handled the technical display issue. It's implemented within &lt;code&gt;llamabot/components/formatters.py&lt;/code&gt; as &lt;code&gt;create_marimo_formatter&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;But to solve the behavioral problem—getting the agent to actually provide contextual text—I had to update the system prompt. I added explicit guidance requiring that every object be preceded (and ideally followed) by explanatory text that connects back to the researcher's goals. The prompt now instructs the agent to create a dictionary with explanatory text strings interleaved with objects, then return it. The formatter processes this dictionary, creating that polished, consultant-like experience where text naturally flows between objects.&lt;/p&gt;
&lt;h2 id="failure-mode-3-agent-decision-making-and-domain-knowledge"&gt;Failure mode 3: Agent decision-making and domain knowledge&lt;/h2&gt;&lt;p&gt;The third category of failures involved the agent's decision-making process and understanding of domain conventions. Through systematic evaluation, I discovered several patterns.&lt;/p&gt;
&lt;p&gt;Here's a concrete example of how the prompt evolved. Initially, I had a simple instruction: "Ask clarifying questions about experiment goals, constraints, and assumptions." But the agent kept jumping straight to calculations. So I added:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Before (early version):&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;Ask clarifying questions about experiment goals, constraints, and assumptions.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;After (refined version):&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;&lt;strong&gt;CRITICAL - BE INQUISITIVE FIRST&lt;/strong&gt;: Before jumping into calculations, you MUST ask probing questions to understand the full context:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What effect size are they expecting or hoping to detect? Why?&lt;/li&gt;
&lt;li&gt;What is the expected variability in their measurements? Do they have pilot data?&lt;/li&gt;
&lt;li&gt;What are their practical constraints (budget, time, sample availability)?&lt;/li&gt;
&lt;li&gt;What are they most worried about with this experiment?&lt;/li&gt;
&lt;li&gt;Have they done similar experiments before? What issues did they encounter?&lt;/li&gt;
&lt;li&gt;What would make this experiment a "success" in their view?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Only after gathering this context&lt;/strong&gt; should you use &lt;code&gt;write_and_execute_code_wrapper&lt;/code&gt; to perform calculations.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;The agent wasn't inquisitive enough.&lt;/strong&gt; Despite the system prompt emphasizing the need to ask questions, the agent would often jump straight to calculations without first understanding the researcher's context, constraints, and goals. I had to add multiple "CRITICAL" reminders in the prompt, explicitly stating that questioning should happen BEFORE calculations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The agent was trying to pass variable names as strings.&lt;/strong&gt; When the agent wanted to analyze a result from a previous tool call, it would try to write a function like &lt;code&gt;def analyze(mtt_power_analysis_result):&lt;/code&gt; and pass &lt;code&gt;{"mtt_power_analysis_result": "mtt_power_analysis_result"}&lt;/code&gt; - which passes the string literal, not the actual dictionary! I had to explicitly teach the agent to write functions with NO parameters that access variables directly from globals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Contradictions in the system prompt.&lt;/strong&gt; There were conflicting instructions about when to use &lt;code&gt;respond_to_user&lt;/code&gt; vs &lt;code&gt;return_object_to_user&lt;/code&gt;. I resolved this by clarifying: use &lt;code&gt;respond_to_user&lt;/code&gt; for text-only responses, and &lt;code&gt;return_object_to_user&lt;/code&gt; when you have Python objects to display.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The agent didn't understand domain-specific visualizations.&lt;/strong&gt; When the test user (a.k.a. me) asked for a "plate map" or "plate layout visualization," the agent would generate something, but it often wasn't what researchers expected.&lt;/p&gt;
&lt;p&gt;A plate map visualization in experimental biology is a very specific thing: a heatmap-style 8×12 grid (rows A-H, columns 1-12) where each well is color-coded by treatment group with a clear legend. Without explicit guidance, the agent would create generic bar charts or scatter plots that didn't match these domain conventions.&lt;/p&gt;
&lt;p&gt;I solved this by adding detailed specifications in the system prompt that describe exactly what plate map visualizations should look like, including the grid structure (rows labeled A-H, columns 1-12 for 96-well plates), color coding requirements (different colors per treatment, with a legend), complete code patterns showing how to create them with matplotlib, common plate formats (96-well, 384-well, 1536-well), and trigger phrases that should generate plate maps.&lt;/p&gt;
&lt;p&gt;This pattern - providing detailed specifications for domain-specific outputs - became a key strategy. When the agent needs to generate something that follows domain conventions, it needs explicit guidance on what those conventions are.&lt;/p&gt;
&lt;p&gt;Addressing these behavioral issues required multiple rounds of iteration. The system prompt didn't reach its final form in one go - as I discovered each issue, I added more explicit guidance, examples, and "CRITICAL" warnings. The prompt grew substantially through iterative refinement, and the agent's behavior improved dramatically with each iteration.&lt;/p&gt;
&lt;p&gt;The process wasn't linear during the early iteration phases - I'd fix one issue, test it, discover another, fix that, and sometimes realize the first fix needed refinement. Working with the Cursor coding agent helped me identify contradictions, explore multiple solution approaches, and get different perspectives before committing to changes. This iterative refinement process is essential when building domain-specific agents: you can't anticipate all the behavioral issues upfront, so you need to be prepared to evolve the prompt based on what you discover through testing.&lt;/p&gt;
&lt;h2 id="conclusion-key-lessons-for-building-domain-specific-agents"&gt;Conclusion: Key lessons for building domain-specific agents&lt;/h2&gt;&lt;p&gt;Building this AI statistics agent for experiment design revealed several important patterns that apply broadly to building domain-specific AI agents and LLM-powered tools:&lt;/p&gt;
&lt;h3 id="1-the-system-prompt-is-the-primary-control-surface"&gt;1. The system prompt is the primary control surface&lt;/h3&gt;&lt;p&gt;The agent's personality, decision-making process, inquisitiveness, and ability to provide contextual explanations are primarily controlled through prompt design rather than code changes. The prompt is where the domain knowledge lives, where the behavioral patterns are encoded, and where the "personality" of the agent is defined. Code provides the infrastructure, but the prompt provides the intelligence.&lt;/p&gt;
&lt;p&gt;If you want to change the agent's behavior, you're often better off modifying the prompt than changing the code. The prompt became a detailed instruction manual that teaches the agent not just &lt;em&gt;what&lt;/em&gt; to do, but &lt;em&gt;how&lt;/em&gt; to think, &lt;em&gt;when&lt;/em&gt; to ask questions, and &lt;em&gt;why&lt;/em&gt; certain patterns matter.&lt;/p&gt;
&lt;h3 id="2-testing-is-essential-but-not-sufficient"&gt;2. Testing is essential, but not sufficient&lt;/h3&gt;&lt;p&gt;Even with extensive testing, you cannot guarantee what will be seen in the real world. Users will ask questions you never thought of, use terminology you didn't anticipate, have edge cases in their data you didn't consider, and interact with the agent in ways that break your assumptions.&lt;/p&gt;
&lt;p&gt;This is why the agent is designed as a &lt;em&gt;consultant&lt;/em&gt; rather than an autonomous decision-maker. A human statistician still needs to review the work. The testing process is essential for building confidence, but you must also design the system with the assumption that it will encounter unexpected situations. This means implementing clear error handling, graceful degradation when things go wrong, explicit boundaries on what the agent can and cannot do, and human oversight for critical decisions.&lt;/p&gt;
&lt;h2 id="final-thoughts"&gt;Final thoughts&lt;/h2&gt;&lt;p&gt;The process of building this agent revealed something I didn't expect: you can't anticipate all the behavioral issues upfront. The prompt grew from ~200 to ~600 lines through iterative discovery—each failure mode required explicit guidance I didn't know I'd need. Building domain-specific agents means being prepared to evolve your approach based on what you discover through testing, not just what you plan in advance.&lt;/p&gt;
&lt;h2 id="try-it-yourself"&gt;Try it yourself&lt;/h2&gt;&lt;p&gt;The experiment design agent is available as a Marimo notebook in the &lt;a href="https://github.com/ericmjl/llamabot"&gt;LlamaBot repository&lt;/a&gt;. You can run it locally with:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;git&lt;span class="w"&gt; &lt;/span&gt;clone&lt;span class="w"&gt; &lt;/span&gt;git@github.com:ericmjl/llamabot.git
&lt;span class="nb"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;llamabot/notebooks
uvx&lt;span class="w"&gt; &lt;/span&gt;marimo&lt;span class="w"&gt; &lt;/span&gt;edit&lt;span class="w"&gt; &lt;/span&gt;--watch&lt;span class="w"&gt; &lt;/span&gt;notebooks/experiment_design_agent.py
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The agent is designed to be inquisitive and consultative—it will ask probing questions about your experiment goals, constraints, and assumptions before providing recommendations. This AI statistics agent can help with power calculations, experimental design critique, sample data table generation, plate map visualizations, and biostatistical consultation for researchers in pharma and biotech.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Limitations&lt;/strong&gt;: This agent is a prototype focused on the experiment design phase. It's not a replacement for human statisticians—it's designed to amplify their knowledge and help researchers think through their designs before data collection. The agent requires human oversight and review, especially for high-stakes decisions. I haven't tested it across all experimental design types, and it may struggle with highly specialized domains or unusual constraints.&lt;/p&gt;
&lt;p&gt;If you're interested in building your own domain-specific agent, I hope the lessons and patterns shared here provide a useful starting point. The code is open source, and I welcome contributions and feedback.&lt;/p&gt;
&lt;p&gt;I'm working on a part 2 of this blog post, where I'll build out the statistical analysis agent—the companion to this experiment design agent. That post will cover how to build an agent that takes collected data and performs statistical analysis, model fitting, and interpretation.&lt;/p&gt;
&lt;p&gt;Thank you for reading this far. If you made it here, you've invested real time and attention in understanding not just what I built, but how and why—and that means a lot! Building this agent has been four months in the making, and sharing those discoveries with others who care about the same problems is what makes the work worthwhile. I'm grateful you came along for the ride!&lt;/p&gt;
</content></entry><entry><title>How to Reference Code Across Repositories with Coding Agents</title><link href="https://ericmjl.github.io/blog/2025/11/17/how-to-reference-code-across-repositories-with-coding-agents/" rel="alternate"/><updated>2025-11-17T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:438621db-6cf5-302d-b316-e2964a8f9e5a</id><content type="html">&lt;p&gt;I used to assume that coding agents like Cursor, GitHub Copilot, and Claude Code only work within a single workspace. This mental model led me to workarounds like copying files, creating complex multi-root workspace configurations, or constantly switching between projects.&lt;/p&gt;
&lt;p&gt;But coding agents can already read and write files from anywhere on your file system, not just the current workspace. The limitation wasn't in the tools; it was in my awareness of what they can do. You don't need to add folders to workspaces, create multi-root workspaces, or jump through configuration hoops. If you know where a repository lives on your disk, you can reference it directly.&lt;/p&gt;
&lt;h2 id="how-to-reference-code-from-other-repositories"&gt;How to reference code from other repositories&lt;/h2&gt;&lt;p&gt;The key is being explicit about file paths. Modern AI coding assistants like Cursor, GitHub Copilot, and Claude Code can access your entire file system, not just the current workspace. You just need to tell them where to look.&lt;/p&gt;
&lt;p&gt;I do most of my writing in an Obsidian vault, which isn't a Git repository; it's just a folder on disk. Sometimes I need to reference code from my LlamaBot repository, or other code repositories in which I am doing development. Instead of copying files or creating complex workspace configurations, I just tell the agent to read directly from the other directory.&lt;/p&gt;
&lt;p&gt;When I need the agent to understand something from LlamaBot, I can say "read the implementation from &lt;code&gt;~/github/llamabot/llamabot/bot/simplebot.py&lt;/code&gt;" and it works immediately. The key is being explicit with the path. You can also search by filename within a directory: "find the notebook named &lt;code&gt;pocketflow_testdrive.py&lt;/code&gt; in &lt;code&gt;~/github/llamabot&lt;/code&gt;". The agent reads the file directly from disk, no workspace configuration needed. You don't need to document paths anywhere; just reference them directly when you need them. That said, if you have commonly accessed paths, documenting them in &lt;code&gt;AGENTS.md&lt;/code&gt; can be helpful for quick reference.&lt;/p&gt;
&lt;p&gt;I used this method while writing my blog post &lt;a href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/"&gt;"How I Replaced 307 Lines of Agent Code with 4 Lines"&lt;/a&gt;. I was drafting the post in my Obsidian vault, but the actual code examples lived in a Marimo notebook within the LlamaBot repository. Rather than copying code snippets or switching workspaces, I had the agent read directly from &lt;code&gt;~/github/llamabot&lt;/code&gt; to pull in the exact implementation details I needed. This let me write about the code while staying in my writing environment, with the agent able to reference the actual source files to ensure accuracy.&lt;/p&gt;
&lt;h2 id="file-system-access-for-ai-coding-assistants-enables-this"&gt;File system access for AI coding assistants enables this&lt;/h2&gt;&lt;p&gt;Coding agents that have file system access can perform read and write operations anywhere they have permission. Tools like Cursor, GitHub Copilot, and Claude Code aren't restricted to the current workspace directory. This works because agents have access to shell tools, the most generic, text-based interface to computers. Shell commands produce text output that agents can read and understand, and they can execute commands anywhere on your system. This means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;You can reference code from any repository on your machine&lt;/li&gt;
&lt;li&gt;You can pull in documentation from other projects&lt;/li&gt;
&lt;li&gt;You can compare implementations across different codebases&lt;/li&gt;
&lt;li&gt;You can reference configuration files from related projects&lt;/li&gt;
&lt;li&gt;You can modify files across multiple repositories when needed&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The only requirement is that you know the path and can tell the agent where to look.&lt;/p&gt;
&lt;h2 id="common-scenarios"&gt;Common scenarios&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;Blogging:&lt;/strong&gt; When writing blog posts about code, reference implementation details from your repositories. The agent can read the actual code to ensure accuracy, pulling in exact examples without copying files or switching workspaces.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Architecture decisions:&lt;/strong&gt; Compare how similar problems are solved across different projects. The agent can read multiple implementations and help you understand trade-offs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Code reuse:&lt;/strong&gt; Before copying code, have the agent check if similar functionality exists elsewhere. It can read files from other repos to find existing solutions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dependency understanding:&lt;/strong&gt; When working with a library you maintain, reference the library's source code directly. The agent can read implementation details to help you use it correctly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cross-repository updates:&lt;/strong&gt; Update related files across multiple repositories simultaneously. For example, update documentation in one repo while modifying the implementation in another, or sync configuration changes across related projects.&lt;/p&gt;
&lt;h2 id="step-by-step-workflow-for-cross-repository-code-access"&gt;Step-by-step workflow for cross-repository code access&lt;/h2&gt;&lt;p&gt;The key trick is being explicit with paths or explicit instructions about how to get to files. Here's how to do it:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;For repositories you already have cloned locally:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Reference the absolute path directly when asking the agent: "read &lt;code&gt;~/github/llamabot/src/llamabot/bot/simple.py&lt;/code&gt;"&lt;/li&gt;
&lt;li&gt;Or search by filename within a directory: "find the notebook named &lt;code&gt;pocketflow_testdrive.py&lt;/code&gt; in &lt;code&gt;~/github/llamabot&lt;/code&gt;"&lt;/li&gt;
&lt;li&gt;The agent reads the file immediately, no workspace configuration needed&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;For repositories you don't have locally:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Tell the agent exactly how to get to the file: "clone the repo &lt;code&gt;owner/repo&lt;/code&gt; into a temporary directory, then find the file at relative path &lt;code&gt;path/to/file.py&lt;/code&gt;"&lt;/li&gt;
&lt;li&gt;You can also specify a specific commit, branch, or tag: "clone the repo &lt;code&gt;owner/repo&lt;/code&gt; at commit &lt;code&gt;abc123&lt;/code&gt; into a temporary directory, then find the file at relative path &lt;code&gt;path/to/file.py&lt;/code&gt;" or "clone the repo &lt;code&gt;owner/repo&lt;/code&gt; and checkout branch &lt;code&gt;feature-branch&lt;/code&gt;, then find the file at relative path &lt;code&gt;path/to/file.py&lt;/code&gt;"&lt;/li&gt;
&lt;li&gt;The agent executes these commands using command line tools like &lt;code&gt;gh&lt;/code&gt; CLI or &lt;code&gt;git&lt;/code&gt;, reads what it needs, and can clean up the temporary clone when done&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;No workspace management. No file copying. No complex configuration. Just explicit paths or explicit instructions. The agent needs clear direction on where to find files, whether that's an absolute path on your system or step-by-step instructions to clone and navigate to a file.&lt;/p&gt;
&lt;h2 id="common-questions-and-limitations"&gt;Common questions and limitations&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;Do I need to configure workspace settings?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;No. Unlike traditional IDE workspace configurations, you don't need to add folders to workspaces or create multi-root setups. Just reference paths directly when you need them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How do I manage paths for many repositories?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;You don't need to document them anywhere. Just reference paths directly when asking the agent to read files. If you find yourself referencing the same paths repeatedly, you can optionally document them in &lt;code&gt;AGENTS.md&lt;/code&gt; for convenience, but it's not required. You can also use MCP server prompts like &lt;code&gt;/remember&lt;/code&gt; (from my &lt;a href="https://github.com/ericmjl/ericmjl-productivity-mcp"&gt;personal productivity MCP server&lt;/a&gt;) to automatically capture frequently-used paths. The &lt;code&gt;/remember&lt;/code&gt; prompt reviews your conversation, identifies important learnings like repository paths, and adds timestamped entries to &lt;code&gt;AGENTS.md&lt;/code&gt; in the appropriate section.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can agents modify files in other repositories?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Yes, but be mindful. While agents can read and write files anywhere on your file system, it's easy to accidentally change files in other repositories. Use this capability deliberately rather than accidentally. Consider using read-only access for cross-repository references unless you specifically need to modify files.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What if I don't have the repository cloned locally?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Have the agent clone it temporarily using command line tools. The agent can use &lt;code&gt;gh&lt;/code&gt; CLI or &lt;code&gt;git&lt;/code&gt; commands to clone repositories into temporary directories, read what it needs, and clean up afterward.&lt;/p&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&lt;p&gt;Thanks to shell tools, coding agents like Cursor, GitHub Copilot, and Claude Code aren't limited by workspace boundaries. They can access your entire file system for both reading and writing, so you can build workflows that span multiple projects without complex tooling.&lt;/p&gt;
&lt;p&gt;The simplicity is the point. You don't need special workspace configurations or multi-root setups. You just need to know where things live and tell the agent where to look. Reference paths directly, or have the agent clone repositories temporarily when needed.&lt;/p&gt;
&lt;p&gt;When you need to reference code from another repository, the agent can read it directly. Just point it to the path. This technique works with any AI coding assistant that has file system access, making it a universal solution for cross-repository code access.&lt;/p&gt;
</content></entry><entry><title>How I Replaced 307 Lines of Agent Code with 4 Lines</title><link href="https://ericmjl.github.io/blog/2025/11/16/how-i-replaced-307-lines-of-agent-code-with-4-lines/" rel="alternate"/><updated>2025-11-16T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:2e208dfe-9d34-3af0-9c37-4a21c5a8528a</id><content type="html">&lt;p&gt;I recently discovered &lt;a href="https://github.com/The-Pocket/PocketFlow?tab=readme-ov-file"&gt;PocketFlow&lt;/a&gt;, a framework for building LLM-enabled programs created by &lt;a href="https://zachary62.github.io/zach_public_material/"&gt;Zachary Huang&lt;/a&gt;. The entire framework is tiny—only 100 lines of code. What caught my attention is that PocketFlow takes a fundamentally different approach to LLM-powered programs, including Anthropic's &lt;a href="https://www.anthropic.com/engineering/building-effective-agents"&gt;workflows and agents&lt;/a&gt;, by structuring them as graphs.&lt;/p&gt;
&lt;p&gt;As someone who used graphs in my thesis work, &lt;a href="https://ericmjl.github.io/Network-Analysis-Made-Simple/"&gt;taught tutorials on applied graph theory&lt;/a&gt;, and &lt;a href="https://github.com/ericmjl/llamabot"&gt;builds my own agent frameworks&lt;/a&gt;, my curiosity was piqued. I wanted to see two things: whether I could learn enough of the framework to build something useful, and whether LlamaBot's abstractions could complement PocketFlow's approach.&lt;/p&gt;
&lt;p&gt;To explore this, I fired up a &lt;a href="https://marimo.io/"&gt;Marimo notebook&lt;/a&gt;. (You can fire it up too by running: &lt;code&gt;uvx marimo edit --sandbox &amp;lt;put URL here to notebook here&amp;gt;&lt;/code&gt;)&lt;/p&gt;
&lt;h2 id="understanding-the-core-nodes-and-flows"&gt;Understanding the Core - Nodes and Flows&lt;/h2&gt;&lt;p&gt;I started by building what I consider a "Hello World" program: a text topic extractor and question generator. This let me familiarize myself with PocketFlow's two core abstractions: &lt;code&gt;Nodes&lt;/code&gt; and &lt;code&gt;Flows&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;A &lt;code&gt;Node&lt;/code&gt; is a unit of execution structured like this:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;SummarizeFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# ...do stuff...&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;stuff_that_gets_passed_to_exec&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_res&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# ...do stuff...&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;stuff_that_gets_passed_to_post&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_res&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# ...do stuff...&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;string_indicator_what_to_do_next&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;There's one more concept to introduce: &lt;code&gt;shared&lt;/code&gt;. In PocketFlow, &lt;code&gt;shared&lt;/code&gt; is like a big workspace that all &lt;code&gt;Node&lt;/code&gt;s can read and write from. Think of it as a kitchen island where chefs and cooks can grab ingredients and leave finished dishes. In computing terms, it's global state that programs can access. In practice, it's simply a dictionary that lives in memory, which any node can manipulate. For example, program &lt;code&gt;memory&lt;/code&gt; might be a key in there, implemented as a list.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;prep -&amp;gt; exec -&amp;gt; post&lt;/code&gt; design within a node is intentional. In theory, you could do everything in one step—there are no hooks that inject stuff between, say, &lt;code&gt;prep&lt;/code&gt; and &lt;code&gt;exec&lt;/code&gt;. In practice, doing everything in one step muddies the program and makes it harder to reason about. I'll show you why later in this post.&lt;/p&gt;
&lt;p&gt;Here's what each step is designed to do:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;prep&lt;/code&gt;&lt;/strong&gt; takes stuff from the &lt;code&gt;shared&lt;/code&gt; dictionary, does any preprocessing, and passes it to &lt;code&gt;exec&lt;/code&gt;. This could include grabbing stuff from memory, interpolating it into a prompt, and returning it for execution with the LLM.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;exec&lt;/code&gt;&lt;/strong&gt; is where the bulk of heavy computation happens. We put API calls to LLM providers (Ollama, OpenAI, Anthropic, etc.) here. What gets returned is passed to the &lt;code&gt;post&lt;/code&gt; method.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;post&lt;/code&gt;&lt;/strong&gt; handles any post-processing. It receives &lt;code&gt;shared&lt;/code&gt;, &lt;code&gt;prep_res&lt;/code&gt; (the result of &lt;code&gt;prep&lt;/code&gt;), and &lt;code&gt;exec_res&lt;/code&gt; (result of &lt;code&gt;exec&lt;/code&gt;). The pattern I've settled on is archiving results in &lt;code&gt;shared&lt;/code&gt;—for example, storing execution results in memory. What gets returned by &lt;code&gt;post&lt;/code&gt; should be a string indicating which downstream path to follow. If nothing specific is needed, it returns &lt;code&gt;default&lt;/code&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A &lt;code&gt;Flow&lt;/code&gt; is declared with a starting &lt;code&gt;Node&lt;/code&gt; and follows the program until completion.&lt;/p&gt;
&lt;p&gt;With this abstraction, multiple LLM-powered abstractions and design patterns can be designed:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://github.com/the-pocket/.github/raw/main/assets/design.png" alt=""&gt;&lt;/p&gt;
&lt;p&gt;(Image from the PocketFlow official documentation.)&lt;/p&gt;
&lt;h2 id="example-1-topic-extractor-and-question-generator"&gt;Example 1 - Topic Extractor and Question Generator&lt;/h2&gt;&lt;p&gt;Here's how I built the two-step/node topic extractor + question generator. First, I declared the nodes:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;ExtractTopics&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;First node: Extract key topics from input text&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;text_to_analyze&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;txt&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;text_to_analyze&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;text_to_analyze&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;text_to_analyze&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;No content to analyze&amp;quot;&lt;/span&gt;

        &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Extract 3-5 key topics from this text. Return only the topics as a comma-separated list:&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;text_to_analyze&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;
        &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SimpleBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;You are a helpful assistant that extracts key topics.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ollama_chat/qwen3:30b&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;topics&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;default&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;GenerateQuestions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Second node: Generate questions based on topics&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;topics&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;topics&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;txt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;txt&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;topics&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;txt&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;topics&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;txt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;topics&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Cannot generate questions without valid topics&amp;quot;&lt;/span&gt;

        &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Given these topics: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;topics&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;and the original text: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;txt&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;Generate 2 interesting questions for each topic.&amp;quot;&lt;/span&gt;
        &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SimpleBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;You are a helpful assistant that generates thought-provoking questions.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ollama_chat/qwen3:30b&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;questions&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;
        &lt;span class="c1"&gt;# No return statement since this is a terminal node.&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Then, I declared the graph:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;extract_topics&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ExtractTopics&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;generate_questions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;GenerateQuestions&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;extract_topics&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;default&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generate_questions&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The magic happens in this line:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;extract_topics&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;default&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;generate_questions&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This tells the flow that once the &lt;code&gt;extract_topics&lt;/code&gt; node emits &lt;code&gt;"default"&lt;/code&gt;, it should proceed to the &lt;code&gt;generate_questions&lt;/code&gt; node. The syntax is compact and looks exactly like an edge specification between two nodes.&lt;/p&gt;
&lt;p&gt;At this point, I deeply appreciate the clarity this approach forces upfront. When thinking about the flow as a graph, I'm forced to think about each node as a function that accepts inputs from shared state and returns a decision about what to do next. That decision can be deterministic (as above) or data-dependent (as we'll see below).&lt;/p&gt;
&lt;p&gt;Since GenAI can be viewed through the lens of automation, &lt;a href="https://ericmjl.github.io/blog/2025/7/13/earn-the-privilege-to-use-automation/"&gt;we should earn the privilege to use it&lt;/a&gt;. Automation requires a well-established process to be most effective. Framing a process in the language of graphs, inputs, and outputs—defining the process as a graph with carefully specified inputs and outputs, just like writing a computer program—is the clearest path to making automation work.&lt;/p&gt;
&lt;p&gt;Running the Flow looks like this:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;shared_topics&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;txt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;txt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;two_node_flow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Flow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;extract_topics&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;two_node_flow&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shared_topics&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;After running, we can inspect the &lt;code&gt;shared_topics&lt;/code&gt; dictionary to see our results:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;txt&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;topics&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="c1"&gt;# added by ExtractTopics&lt;/span&gt;
    &lt;span class="s2"&gt;&amp;quot;questions&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# added by GenerateQuestions&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;One thing missing from PocketFlow is the ability to visualize the graph directly. Since the codebase was new to me, I sent a Cursor agent in the background to research and propose a solution. It came back with &lt;a href="https://github.com/ericmjl/llamabot/pull/279"&gt;this PR&lt;/a&gt;. Impressive!&lt;/p&gt;
&lt;p&gt;The Mermaid diagram for this workflow is:&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph LR
N1["ExtractTopics"]
N2["GenerateQuestions"]
N1 --&gt; N2
style N1 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N2 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;

&lt;/pre&gt;&lt;h2 id="example-2-building-an-agent"&gt;Example 2 - Building an Agent&lt;/h2&gt;&lt;p&gt;Now, what if we want to build an agent?&lt;/p&gt;
&lt;p&gt;I'm going to work backwards here. My "hello world" test for agentic systems is making them tell me today's date. This works because an LLM will always hallucinate a date on its own, and that hallucination may or may not be correct. An agent that works properly should call a tool to get the actual date. The agent's graph should look like this:&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph LR
    N1["Decide"]
    N2["TodayDate"]
    N3["RespondToUser"]
    N1 --&gt;|"today_date"| N2
    N2 --&gt;|"decide"| N1
    N1 --&gt;|"respond_to_user"| N3
    style N1 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
    style N2 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
    style N3 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;

&lt;/pre&gt;&lt;p&gt;I consider this a "Hello World" agent because a failing agent will skip straight to &lt;code&gt;respond_to_user&lt;/code&gt; when asked for today's date, without first calling &lt;code&gt;today_date&lt;/code&gt; to get the actual information.&lt;/p&gt;
&lt;p&gt;To build this agent, I need three nodes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;Decide&lt;/code&gt;&lt;/strong&gt;: Uses an LLM to decide which tool to call next, given the prompt&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;TodayDate&lt;/code&gt;&lt;/strong&gt;: Executes without LLMs and returns today's date in the current timezone&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;RespondToUser&lt;/code&gt;&lt;/strong&gt;: Responds to the user with the appropriate context&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here's how I wrote them. First, the &lt;code&gt;Decide&lt;/code&gt; node:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;llamabot.components.tools&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;respond_to_user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;search_internet&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;today_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;pydantic&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;typing&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;

&lt;span class="n"&gt;search_internet&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;search_internet&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;respond_to_user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;today_date&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;ToolChoice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="vm"&gt;__name__&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;The name of the tool to use&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="nd"&gt;@lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;system&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;decision_bot_system_prompt&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Given the chat history, pick for me one or more tools to execute&lt;/span&gt;
&lt;span class="sd"&gt;    in order to satisfy the user&amp;#39;s query.&lt;/span&gt;

&lt;span class="sd"&gt;    Give me just the tool name to pick.&lt;/span&gt;
&lt;span class="sd"&gt;    Use the tools judiciously to help answer the user&amp;#39;s query.&lt;/span&gt;
&lt;span class="sd"&gt;    Query is always related to one of the tools.&lt;/span&gt;
&lt;span class="sd"&gt;    Use respond_to_user if you have enough information to answer the original query.&lt;/span&gt;
&lt;span class="sd"&gt;    &amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;Decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Query: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;query&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;pydantic_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ToolChoice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;decision_bot_system_prompt&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nb"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Chosen Tool: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The key thing to note is that we inject the available tools into the system prompt of the tool-selecting agent.&lt;/p&gt;
&lt;p&gt;Next, the &lt;code&gt;TodayDate&lt;/code&gt; node:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;TodayDate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;today_date&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Today&amp;#39;s date: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And finally, the &lt;code&gt;RespondToUser&lt;/code&gt; node:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;RespondToUser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;The response to the user.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="s2"&gt;&amp;quot;You are a helpful assistant.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ollama_chat/gemma3n:latest&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;pydantic_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Finally, we set up the graph:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# Set up the graph&lt;/span&gt;
&lt;span class="n"&gt;today__date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;TodayDate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;respond__to__user&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;RespondToUser&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Decide&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;shared&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;query&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;What is the date today?&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;today_date&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;today__date&lt;/span&gt;
&lt;span class="n"&gt;today__date&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;decide&lt;/span&gt;
&lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;respond_to_user&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;respond__to__user&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;I used &lt;code&gt;__&lt;/code&gt; in the node names to avoid clashing with the original functions.&lt;/p&gt;
&lt;p&gt;Then we run it:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;flow2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Flow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;flow2&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Take my word for it (or check out the notebook yourself)—it reliably gives me today's date.&lt;/p&gt;
&lt;h2 id="example-3-agent-with-shell-commands"&gt;Example 3 - Agent with Shell Commands&lt;/h2&gt;&lt;p&gt;To push things further, I tried a tool that needs arguments. A good "hello world" for this is executing shell commands in response to questions like, "What's in this folder?"&lt;/p&gt;
&lt;p&gt;For this, I created a second version of the &lt;code&gt;Decide&lt;/code&gt; node called &lt;code&gt;Decide2&lt;/code&gt;, where I instantiate and execute the &lt;code&gt;ToolChoice&lt;/code&gt; and tool selection &lt;code&gt;StructuredBot&lt;/code&gt; within &lt;code&gt;exec&lt;/code&gt;:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;Decide2&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;sysprompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;decision_bot_system_prompt&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="nb"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sysprompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;ToolChoice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="vm"&gt;__name__&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;tools&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]]]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;The name of the tool to use&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;justification&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Why this tool was chosen.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;pydantic_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ToolChoice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;decision_bot_system_prompt&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;query&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Query: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;query&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Chosen Tool: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;I then created a &lt;code&gt;ShellCommand&lt;/code&gt; node that uses the same pattern—leveraging &lt;code&gt;StructuredBot&lt;/code&gt; for structured generation to constrain the LLM's output to exactly what I need:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;ShellCommand&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;Cmd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;The shell command to execute&amp;quot;&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;You are an expert at writing shell commands. For the chat trace that you will be given, write a shell command that accomplishes the user&amp;#39;s request. Only output the command, nothing else.&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;pydantic_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Cmd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ollama_chat/gemma3n:latest&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="nb"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;execute_shell_command&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Output: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exec_result&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nb"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Finally, we set up the graph:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;_&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# Set up the graph&lt;/span&gt;
    &lt;span class="n"&gt;today_date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;TodayDate&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;respond_to_user&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;RespondToUser&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Decide2&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;shell_command&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ShellCommand&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;today_date&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;today_date&lt;/span&gt;
    &lt;span class="n"&gt;today_date&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;decide&lt;/span&gt;
    &lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;execute_shell_command&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;shell_command&lt;/span&gt;
    &lt;span class="n"&gt;shell_command&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;decide&lt;/span&gt;
    &lt;span class="n"&gt;decide&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;respond_to_user&amp;quot;&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;respond_to_user&lt;/span&gt;

    &lt;span class="n"&gt;flow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Flow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;flow&lt;/span&gt;


&lt;span class="n"&gt;flow3&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The graph would look like this:&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph LR
N1["Decide2"]
N2["TodayDate"]
N3["ShellCommand"]
N4["RespondToUser"]
N1 --&gt;|"today_date"| N2
N2 --&gt;|"decide"| N1
N1 --&gt;|"execute_shell_command"| N3
N3 --&gt;|"decide"| N1
N1 --&gt;|"respond_to_user"| N4
style N1 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N2 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N3 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N4 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;

&lt;/pre&gt;&lt;p&gt;I wrapped it in a &lt;code&gt;_()&lt;/code&gt; function to protect the globally scoped variables in the Marimo notebook. Note that I included &lt;code&gt;today_date&lt;/code&gt; as well, just to "pollute" the namespace and make it more challenging when asking shell-related questions. When we interact with the agent:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="n"&gt;shared3&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;shared3&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;query&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;What in my current working directory?&amp;quot;&lt;/span&gt;
&lt;span class="n"&gt;shared3&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="n"&gt;shared3&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;tools&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;respond_to_user&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;today_date&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;execute_shell_command&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;flow3&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shared3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;It calls on &lt;code&gt;shell_command&lt;/code&gt;, and gives me back this response:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;Okay, here&amp;#39;s a list of the files and directories in your current working directory:

*   **Directories:**
    *   `__marimo__`
    *   `.` (current directory)
    *   `..` (parent directory)

*   **Files:**
    *   `agentbot_build.py`
    *   `agents.py`
    *   `chatbot_as_agent.py`
    *   `conversation-threads.py`
    *   `data.csv`
    *   `ic50_data_with_confounders.csv`
    *   `intro.py`
    *   `lancedb_docstore.py`
    *   `pocketflow_testdrive.py`
    *   `react-agentbot-demo.py`
    *   `README.md`
    *   `toolbot_chatdata.py`
    *   `tools.py`

That&amp;#39;s a total of 17 files and directories. Let me know if you&amp;#39;d like more details about any of them!
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;I was also able to ask "Hey, what files have been modified today?" and the agent successfully executed the appropriate shell command.&lt;/p&gt;
&lt;p&gt;Effectively, this pattern is nothing more than a coordinating agent/LLM delegating work to specialized tools.&lt;/p&gt;
&lt;h2 id="rewriting-agentbot-with-pocketflow"&gt;Rewriting AgentBot with PocketFlow&lt;/h2&gt;&lt;p&gt;Finally, I decided to take what I'd learned and redo the &lt;code&gt;AgentBot&lt;/code&gt; implementation in LlamaBot. My previous implementation (version 0.16.3) was messy—the &lt;code&gt;__call__&lt;/code&gt; method alone was 307 lines with a while-loop, maximum tries, ThreadPoolExecutor for parallel tool execution, tool call caching, and extensive metadata tracking. PocketFlow had a better abstraction for the agentic loop: a &lt;code&gt;Flow&lt;/code&gt; state machine following edges on a graph. I thought I could redesign &lt;code&gt;AgentBot&lt;/code&gt; to take advantage of this pattern.&lt;/p&gt;
&lt;p&gt;The rewrite involved some really interesting patterns. I completely replaced the ReAct (Reasoning and Acting) loop with PocketFlow's graph-based tool orchestration. This shifts from an iterative loop-based approach to a declarative graph-based one, where tool execution flows through a directed graph rather than a sequential loop.&lt;/p&gt;
&lt;p&gt;The implementation centers on three key abstractions:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. The &lt;code&gt;@nodeify&lt;/code&gt; decorator&lt;/strong&gt; transforms any callable function into a PocketFlow Node. It wraps functions with PocketFlow's Node interface, implementing the required &lt;code&gt;prep&lt;/code&gt;, &lt;code&gt;exec&lt;/code&gt;, and &lt;code&gt;post&lt;/code&gt; methods. The tricky part is that &lt;code&gt;@nodeify&lt;/code&gt; needs to preserve access to the underlying function's metadata—particularly the &lt;code&gt;json_schema&lt;/code&gt; attribute added by the &lt;code&gt;@tool&lt;/code&gt; decorator—through attribute proxying, so ToolBot can discover and use tools even after they've been wrapped as nodes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. The &lt;code&gt;DecideNode&lt;/code&gt;&lt;/strong&gt; encapsulates the decision-making logic. This node uses ToolBot internally to analyze the conversation history stored in shared state and select which tool to execute next. It expects a shared state dictionary with a &lt;code&gt;"memory"&lt;/code&gt; key containing the conversation history as a list of strings. When executed, it calls ToolBot with this memory, extracts the first tool call from ToolBot's response, parses the JSON-formatted arguments, and stores them in &lt;code&gt;shared["func_call"]&lt;/code&gt; for the next node. The node then returns the tool name as a routing action, which PocketFlow uses to navigate the graph.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Flow graph construction&lt;/strong&gt; happens at initialization time. AgentBot automatically wraps all provided tools (plus default tools like &lt;code&gt;today_date&lt;/code&gt; and &lt;code&gt;respond_to_user&lt;/code&gt;) with both &lt;code&gt;@tool&lt;/code&gt; and &lt;code&gt;@nodeify&lt;/code&gt; decorators, then builds bidirectional connections: from the decide node to each tool node (using the tool's function name as the action), and from each tool node back to the decide node (except for terminal tools like &lt;code&gt;respond_to_user&lt;/code&gt; that have &lt;code&gt;loopback_name=None&lt;/code&gt;). This creates a graph where execution can flow from decision to tool and back to decision, enabling multi-step reasoning.&lt;/p&gt;
&lt;p&gt;A few technical requirements make this work: tools need type annotations (for JSON schema generation), the shared state needs a &lt;code&gt;"memory"&lt;/code&gt; list for conversation history, and tool arguments are passed through &lt;code&gt;shared["func_call"]&lt;/code&gt;. The DecideNode selects one tool at a time, and tools are stateless—they get fresh arguments each call and communicate through memory.&lt;/p&gt;
&lt;p&gt;What's remarkable about this implementation is how compact it is. The &lt;code&gt;@nodeify&lt;/code&gt; decorator is just 100 lines, and most of that is documentation. The core logic is elegant:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;nodeify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;loopback_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;FuncNode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="fm"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="nb"&gt;super&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="fm"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;loopback_name&lt;/span&gt;
                &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;func&lt;/span&gt;

            &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;prep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;

            &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="n"&gt;func_call&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;func_call&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{})&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;func_call&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

            &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prep_result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exec_res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;exec_res&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt;

            &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="fm"&gt;__getattr__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="c1"&gt;# Proxy to original function for json_schema access&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;func&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="ne"&gt;AttributeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;getattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;FuncNode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;func&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The entire AgentBot class is similarly compact—about 100 lines total. Compare this to the previous implementation where the &lt;code&gt;__call__&lt;/code&gt; method alone was 307 lines, with complex while loop logic, tool call caching, parallel execution via ThreadPoolExecutor, and extensive state management:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# Old implementation (v0.16.3): 307-line __call__ method&lt;/span&gt;
&lt;span class="c1"&gt;# Plus 50-line caching wrapper, 21-line execution helper&lt;/span&gt;
&lt;span class="c1"&gt;# Total: 378 lines of orchestration code&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;iteration&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nb"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_iterations&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Call model with tools&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;completion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;raw_messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;tool_choice&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;auto&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;tool_calls&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;extract_tool_calls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Execute tools in parallel with caching&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;futures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;submit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_execute_tool_with_cache&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;tool_calls&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;future&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;as_completed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;futures&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="c1"&gt;# Handle results, update messages, manage cache,&lt;/span&gt;
                &lt;span class="c1"&gt;# track metadata, handle errors...&lt;/span&gt;
                &lt;span class="o"&gt;...&lt;/span&gt;
        &lt;span class="k"&gt;continue&lt;/span&gt;

    &lt;span class="c1"&gt;# Handle finalization, memory updates, logging, metrics...&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The new implementation replaces all of that with a simple graph construction:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="c1"&gt;# New implementation: ~100 lines total, declarative graph&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;AgentBot&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="fm"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;decide_node&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;gpt-4.1&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# ... validation and setup ...&lt;/span&gt;

        &lt;span class="c1"&gt;# Build PocketFlow graph: connect tools to decide node&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;all_tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;tool_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="vm"&gt;__name__&lt;/span&gt;
            &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decide_node&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;tool_name&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;tool_node&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decide_node&lt;/span&gt;

        &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Flow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decide_node&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="fm"&gt;__call__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;memory&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flow&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shared&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;result&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Full implementation includes validation, tool wrapping, and state management—about 100 lines total vs 307+ for the old &lt;code&gt;__call__&lt;/code&gt; method alone.&lt;/p&gt;
&lt;h3 id="the-magic-of-building-an-agent-in-just-4-lines"&gt;The Magic of Building an Agent in Just 4 Lines&lt;/h3&gt;&lt;p&gt;The most remarkable part of this implementation is how the entire agent graph is constructed. Look at these four lines carefully:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;all_tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decide_node&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;tool_name&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;tool_node&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;tool_node&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decide_node&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;This is it.&lt;/strong&gt; This is the entire graph construction that turns a collection of tools into a working agent. Let me break down what's happening:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Line 1&lt;/strong&gt;: Loop through each tool&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Line 2&lt;/strong&gt;: Connect the decide node to the tool node—when the LLM chooses this tool, execution flows to it&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Line 3&lt;/strong&gt;: Check if this tool should loop back (terminal tools like &lt;code&gt;respond_to_user&lt;/code&gt; have &lt;code&gt;loopback_name=None&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Line 4&lt;/strong&gt;: Connect the tool back to the decide node—after execution, control returns to decision-making&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That's the entire agent architecture. Four lines. The &lt;code&gt;- "action" &amp;gt;&amp;gt;&lt;/code&gt; syntax creates directed edges in the graph, and PocketFlow handles all the state management, routing, and execution orchestration. Compare this to the 307-line &lt;code&gt;__call__&lt;/code&gt; method in the previous implementation (version 0.16.3) with its complex loop-based logic, thread pools, state tracking, and termination conditions.&lt;/p&gt;
&lt;p&gt;This is what I mean by "graph-based thinking" being clearer—the entire execution flow is explicit and declarative. You can see at a glance how decisions flow to tools and back to decisions, enabling multi-step reasoning.&lt;/p&gt;
&lt;p&gt;The difference is striking. The old implementation required manual loop management, explicit state tracking, parallel execution coordination, and complex termination logic. The new implementation declares the graph structure once, and PocketFlow handles all the execution details.&lt;/p&gt;
&lt;p&gt;This graph-based approach provides several advantages. The flow graph is constructed once at initialization, making the execution path explicit and visualizable—you can render the agent's decision flow as a Mermaid diagram using the visualization feature I added to LlamaBot. The separation of concerns is clearer: decision-making lives in &lt;code&gt;DecideNode&lt;/code&gt;, tool execution in wrapped function nodes, and orchestration in PocketFlow's flow engine. The implementation is also more modular—you can swap out the decision node or customize tool wrapping behavior without rewriting the core agent logic. Finally, by leveraging PocketFlow's graph execution model, we gain access to its execution capabilities and potential future extensions for parallel execution or conditional routing.&lt;/p&gt;
&lt;h2 id="visualizing-different-agent-architectures"&gt;Visualizing Different Agent Architectures&lt;/h2&gt;&lt;p&gt;One really cool feature I added to LlamaBot is the ability to visualize any agent's graph structure using Mermaid diagrams. The &lt;code&gt;AgentBot._display_()&lt;/code&gt; method automatically renders the flow graph, making it easy to see how different tool configurations create different architectures.&lt;/p&gt;
&lt;p&gt;Here's a simple agent with just two tools:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;llamabot&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AgentBot&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;llamabot.components.tools&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;llamabot.components.pocketflow&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;nodeify&lt;/span&gt;

&lt;span class="nd"&gt;@nodeify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@tool&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;search_web&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Search the web for information.&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;web_search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;search_web&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_display_&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# Renders Mermaid diagram in Marimo&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The resulting graph shows the decision node connected to &lt;code&gt;today_date&lt;/code&gt;, &lt;code&gt;search_web&lt;/code&gt;, and &lt;code&gt;respond_to_user&lt;/code&gt;:&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph LR
N1["DecideNode"]
N2["today_date"]
N3["search_web"]
N4["respond_to_user"]
N1 --&gt;|"today_date"| N2
N2 --&gt;|"decide"| N1
N1 --&gt;|"search_web"| N3
N3 --&gt;|"decide"| N1
N1 --&gt;|"respond_to_user"| N4
style N1 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N2 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N3 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N4 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
&lt;/pre&gt;&lt;p&gt;Add more tools, and the graph automatically expands. Here's an agent with code execution and file operations:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nd"&gt;@nodeify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@tool&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;write_and_execute_script&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;dependencies_str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;python_version&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;&amp;gt;=3.11&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Write and execute a Python script in a secure Docker sandbox.&lt;/span&gt;

&lt;span class="sd"&gt;    :param code: The Python code to execute&lt;/span&gt;
&lt;span class="sd"&gt;    :param dependencies_str: Comma-separated pip dependencies&lt;/span&gt;
&lt;span class="sd"&gt;    :param python_version: Python version requirement&lt;/span&gt;
&lt;span class="sd"&gt;    :return: Dictionary with stdout, stderr, and status&lt;/span&gt;
&lt;span class="sd"&gt;    &amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;
    &lt;span class="c1"&gt;# Uses ScriptExecutor to run code in isolated Docker container&lt;/span&gt;
    &lt;span class="n"&gt;executor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ScriptExecutor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;executor&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;run_script&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;script_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;stdout&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;stdout&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;stderr&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;stderr&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;status&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;status&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nd"&gt;@nodeify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@tool&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;read_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filepath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Read and return file contents.&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filepath&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;search_web&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;write_and_execute_script&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;read_file&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The graph now shows six tool nodes all connected bidirectionally to the decision node (except terminal tools):&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph LR
N1["DecideNode"]
N2["today_date"]
N3["search_web"]
N4["write_and_execute_script"]
N5["read_file"]
N6["respond_to_user"]
N7["return_object_to_user"]
N1 --&gt;|"today_date"| N2
N2 --&gt;|"decide"| N1
N1 --&gt;|"search_web"| N3
N3 --&gt;|"decide"| N1
N1 --&gt;|"write_and_execute_script"| N4
N4 --&gt;|"decide"| N1
N1 --&gt;|"read_file"| N5
N5 --&gt;|"decide"| N1
N1 --&gt;|"respond_to_user"| N6
N1 --&gt;|"return_object_to_user"| N7
style N1 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N2 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N3 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N4 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N5 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N6 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N7 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
&lt;/pre&gt;&lt;p&gt;What I love about this is how the graph makes it immediately obvious what capabilities an agent has. You can see at a glance which tools are available, understand the control flow, and reason about how the agent will behave. The visualization transforms the abstract "agent with tools" into a concrete, inspectable structure.&lt;/p&gt;
&lt;p&gt;Here's a real-world example—an experiment design agent I built for critiquing statistical experiment designs:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nd"&gt;@nodeify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;loopback_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;decide&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@tool&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;critique_experiment_design&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;design&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Critique an experiment design and identify potential flaws,&lt;/span&gt;
&lt;span class="sd"&gt;    biases, or weaknesses.&lt;/span&gt;

&lt;span class="sd"&gt;    :param design: Description of the proposed experiment design&lt;/span&gt;
&lt;span class="sd"&gt;    :return: Critique with identified issues and suggestions&lt;/span&gt;
&lt;span class="sd"&gt;    &amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SimpleBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;experiment_design_critique_sysprompt&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;design&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;critique_experiment_design&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;write_and_execute_code&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;globals&lt;/span&gt;&lt;span class="p"&gt;())]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This agent has a specialized domain focus. The graph shows all its capabilities, including the default tools that every &lt;code&gt;AgentBot&lt;/code&gt; gets automatically:&lt;/p&gt;
&lt;pre class="mermaid"&gt;
graph LR
N1["DecideNode"]
N2["today_date"]
N3["critique_experiment_design"]
N4["write_and_execute_code"]
N5["respond_to_user"]
N6["return_object_to_user"]
N1 --&gt;|"today_date"| N2
N2 --&gt;|"decide"| N1
N1 --&gt;|"critique_experiment_design"| N3
N3 --&gt;|"decide"| N1
N1 --&gt;|"write_and_execute_code"| N4
N4 --&gt;|"decide"| N1
N1 --&gt;|"respond_to_user"| N5
N1 --&gt;|"return_object_to_user"| N6
style N1 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N2 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N3 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N4 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N5 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
style N6 fill:#e1f5ff,stroke:#01579b,stroke-width:2px;
&lt;/pre&gt;&lt;p&gt;Notice that &lt;code&gt;today_date&lt;/code&gt;, &lt;code&gt;respond_to_user&lt;/code&gt;, and &lt;code&gt;return_object_to_user&lt;/code&gt; are included by default in every &lt;code&gt;AgentBot&lt;/code&gt;. The graph immediately tells you this agent can critique designs, execute code to analyze data, and return Python objects directly to the user—but it's not a general-purpose assistant. It's specialized for experiment design evaluation. The visual structure encodes the agent's purpose.&lt;/p&gt;
&lt;p&gt;This is only possible because of the graph-based architecture. With the old loop-based implementation, there was no clean way to visualize the execution flow—it was hidden inside imperative control logic.&lt;/p&gt;
&lt;h2 id="what-i-learned"&gt;What I Learned&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;Externalize memory as shared state&lt;/strong&gt;. Memory lives in the &lt;code&gt;shared&lt;/code&gt; dictionary that all nodes can access, rather than being intrinsic to each bot. We just feed memory context in each time a node executes. This has good economics—if you have prompt caching on the API provider's side, simply appending to an ever-growing memory is a great way to take advantage of pre-computed neural network outputs from previous runs. I used to think of memory as &lt;em&gt;intrinsic&lt;/em&gt; to a bot, but I've changed my mind: allowing multiple bots to share access to the same memory is a useful simplification, even if it's not suitable for every circumstance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The &lt;code&gt;prep -&amp;gt; exec -&amp;gt; post&lt;/code&gt; pattern&lt;/strong&gt; is worth adhering to. I found myself appending to memory in &lt;code&gt;post&lt;/code&gt; after doing the &lt;code&gt;exec&lt;/code&gt;ution. &lt;code&gt;prep&lt;/code&gt; turns out to be useful for preprocessing user inputs or manipulating memory as needed. The overall effect is that it's much easier to &lt;strong&gt;unit test or set up evals&lt;/strong&gt; for individual nodes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PocketFlow's graph abstraction brings clarity&lt;/strong&gt;. The analogy of LLM agents (&lt;code&gt;Node&lt;/code&gt;s) as chefs/cooks accessing a kitchen island's worth of things (&lt;code&gt;shared&lt;/code&gt;) is a powerful contrast to my previous loop-based approach in LlamaBot, where I manually tracked state, managed iterations, and coordinated tool execution. This insight is exactly why I rewrote AgentBot to use this graph-based architecture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A wide variety of LLM-powered architectures&lt;/strong&gt; can be built with just &lt;code&gt;Node&lt;/code&gt;s and &lt;code&gt;Flow&lt;/code&gt;s. Most LLM applications I've built—whether for myself or for others—have not been "agentic" but more like "workflows." These are what some might consider boring. Yet they are high ROI precisely because they take repetitive and boring work out of our hands! PocketFlow gives us a way to express flows as graphs, effectively state machines whose actions are either fully deterministic or determined by an LLM's choice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PocketFlow is minimalistic&lt;/strong&gt;, which offloads a lot of heavy lifting when working with LLMs. The flexibility is both a strength and a weakness: great for power users, but potentially intimidating for newcomers. I found it easiest to rely heavily on &lt;code&gt;StructuredBot&lt;/code&gt; to output decisions made by the LLM. Structured generation is, generally speaking, the most useful abstraction in the LLM world that I keep turning back to.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The agent pattern is everywhere once you recognize it&lt;/strong&gt;. While writing this post, I realized that the agentic coding IDEs we've gotten used to—tools like Cursor, GitHub Copilot, and others—follow the exact same pattern I've been describing. They have a decision node that analyzes your code and context, tool nodes for reading files, searching codebases, editing code, and responding to you. The flow is the same: decide what to do, execute a tool, update context, decide again. Understanding this pattern in PocketFlow helped me see it operating in the tools I use every day. The abstraction is the mental model that makes sense of how modern AI-powered tools work.&lt;/p&gt;
&lt;p&gt;The biggest lesson? &lt;strong&gt;Thinking in graphs transforms how you build LLM programs&lt;/strong&gt;. The shift from imperative loops to declarative graphs means you declare &lt;em&gt;what&lt;/em&gt; should happen instead of specifying &lt;em&gt;how&lt;/em&gt; to execute step-by-step. This brings clarity, modularity, and makes your execution flow explicit. Whether you're building simple workflows or complex agents, representing them as graphs forces you to think clearly about state, decisions, and flow. That mental model shift has changed how I approach every LLM application I build.&lt;/p&gt;
</content></entry><entry><title>Safe ways to let your coding agent work autonomously</title><link href="https://ericmjl.github.io/blog/2025/11/8/safe-ways-to-let-your-coding-agent-work-autonomously/" rel="alternate"/><updated>2025-11-08T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:0da86835-0adf-31d4-be7c-70ec5f74e11d</id><content type="html">&lt;p&gt;Coding agents promise to unlock significant productivity gains by working autonomously in the background—gathering context, running tests, searching documentation, and making progress on tasks without constant human intervention. The more autonomous they become, the more value they deliver. Yet this autonomy creates a fundamental tension: we need agents to act independently to realize their potential, but we must prevent them from taking irreversible actions we don't want.&lt;/p&gt;
&lt;p&gt;This tension became painfully clear when I asked Comet, an agentic browser, "how to archive repo" in the same casual way I'd ask Google. The agent interpreted this as a direct command and archived my LlamaBot repository. What I wanted was information; what I got was an unintended action with real consequences.&lt;/p&gt;
&lt;p&gt;The problem isn't unique to Comet. Any coding agent with sufficient autonomy can make destructive changes: deleting files, force-pushing to main, committing broken code, or modifying critical configurations. We need safeguards that allow agents to work freely on safe operations while blocking potentially harmful actions. The solution lies in configuring your development environment with intelligent boundaries—auto-approving read-only commands while requiring explicit approval for anything that modifies state.&lt;/p&gt;
&lt;h2 id="auto-approve-safe-command-line-commands"&gt;Auto-approve safe command line commands&lt;/h2&gt;&lt;p&gt;The foundation of autonomous coding agent operation is allowing certain command line commands to run without manual approval. Commands like &lt;code&gt;grep&lt;/code&gt;/&lt;code&gt;ripgrep&lt;/code&gt;, &lt;code&gt;find&lt;/code&gt;/&lt;code&gt;fd&lt;/code&gt;, &lt;code&gt;pixi run pytest...&lt;/code&gt;, and similar read-only or context-gathering operations enable LLM agents to autonomously understand codebases and test suites. For CLI tools that interact with external services, I also auto-approve &lt;code&gt;gh pr view&lt;/code&gt;, which allows the agent to gather context from GitHub pull requests while working in the background.&lt;/p&gt;
&lt;p&gt;The critical rule: &lt;strong&gt;only auto-accept commands that are non-destructive&lt;/strong&gt;. Never auto-approve &lt;code&gt;git commit&lt;/code&gt;, &lt;code&gt;git push&lt;/code&gt;, &lt;code&gt;rm&lt;/code&gt;, or other filesystem, git, or state-modifying changes. This creates a safe boundary where agents can explore and learn, but cannot make irreversible changes without your explicit approval.&lt;/p&gt;
&lt;p&gt;Here's my mental model for categorizing commands:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Safe to auto-approve:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Read operations: &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;find&lt;/code&gt;, &lt;code&gt;cat&lt;/code&gt;, &lt;code&gt;head&lt;/code&gt;, &lt;code&gt;tail&lt;/code&gt;, &lt;code&gt;less&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Code analysis: &lt;code&gt;pytest&lt;/code&gt; (read-only test runs), &lt;code&gt;mypy&lt;/code&gt;, &lt;code&gt;ruff check&lt;/code&gt; (without &lt;code&gt;--fix&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Context gathering: &lt;code&gt;gh pr view&lt;/code&gt;, &lt;code&gt;gh issue view&lt;/code&gt;, &lt;code&gt;git log&lt;/code&gt;, &lt;code&gt;git diff&lt;/code&gt;, &lt;code&gt;git show&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Package managers (read-only): &lt;code&gt;pip list&lt;/code&gt;, &lt;code&gt;npm list&lt;/code&gt;, &lt;code&gt;cargo tree&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Documentation build: &lt;code&gt;mkdocs serve&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Never auto-approve:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;File system mutations: &lt;code&gt;rm&lt;/code&gt;, &lt;code&gt;mv&lt;/code&gt;, &lt;code&gt;cp&lt;/code&gt;, &lt;code&gt;mkdir&lt;/code&gt;, &lt;code&gt;touch&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Git writes: &lt;code&gt;git commit&lt;/code&gt;, &lt;code&gt;git push&lt;/code&gt;, &lt;code&gt;git reset&lt;/code&gt;, &lt;code&gt;git checkout -b&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Package installs: &lt;code&gt;pixi add&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The edge cases are where it gets interesting. I auto-approve &lt;code&gt;pytest&lt;/code&gt; because test runs are read-only, but I require approval for any command that modifies files, even if it's technically reversible. The key distinction is whether a command changes state: &lt;code&gt;git status&lt;/code&gt; and &lt;code&gt;git diff&lt;/code&gt; are safe because they're pure reads, while &lt;code&gt;git commit&lt;/code&gt; and &lt;code&gt;git push&lt;/code&gt; modify repository state and require explicit approval. &lt;code&gt;git add&lt;/code&gt; is a bit of a gray area, but I am ok with auto-approving it since it's technically reversible, and because coding agents are often much faster than I could be at selectively adding files to the staging area.&lt;/p&gt;
&lt;h2 id="enable-automatic-web-search"&gt;Enable automatic web search&lt;/h2&gt;&lt;p&gt;For Cursor and Claude Code, automatic web &lt;em&gt;search&lt;/em&gt; without approval requests is another powerful capability. I have web search auto-approved on my machine, which allows agents to look up documentation, error messages, and solutions independently. This is particularly valuable when agents encounter unfamiliar error messages or need to check current API documentation that may have changed since the model's training cutoff.&lt;/p&gt;
&lt;p&gt;However, I monitor outputs for prompt poisoning, since internet-based prompt poisoning is a known attack vector for AI systems. The risk is that malicious content from web searches could influence the agent's behavior in subsequent actions. I've found this risk manageable for coding tasks, but I'm more cautious with agents that have broader system access or handle sensitive data.&lt;/p&gt;
&lt;h2 id="know-your-emergency-stop-shortcuts"&gt;Know your emergency stop shortcuts&lt;/h2&gt;&lt;p&gt;Every coding agent platform provides keyboard shortcuts to cancel actions in progress. These are essential when you notice an agent looping, going down an unproductive path, or making changes you don't want:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cursor: &lt;code&gt;Ctrl+C&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;VSCode + GitHub Copilot: &lt;code&gt;Cmd+Esc&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Claude Code: &lt;code&gt;Esc&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you're monitoring the agent's activity, these shortcuts let you intervene immediately when something goes wrong.&lt;/p&gt;
&lt;h2 id="correct-agent-behavior-in-real-time"&gt;Correct agent behavior in real-time&lt;/h2&gt;&lt;p&gt;When you catch an agent doing something undesirable, stop it immediately, then redirect it. I instruct agents to record corrections in &lt;code&gt;AGENTS.md&lt;/code&gt; and continue with the updated guidance. An example prompt:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;No, I don&amp;#39;t want you to do &amp;lt;thing&amp;gt;. Instead, you should do &amp;lt;a different thing&amp;gt;. Record this in AGENTS.md, and then continue what you were doing.
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This approach creates a persistent record of preferences that improves future agent behavior. The &lt;code&gt;AGENTS.md&lt;/code&gt; file becomes a living document of your development standards and preferences, which agents can reference in future sessions. I've implemented this pattern in my &lt;a href="https://github.com/ericmjl/ericmjl-productivity-mcp"&gt;personal productivity MCP server&lt;/a&gt;, which provides a standardized way to store and retrieve these preferences across different agent platforms.&lt;/p&gt;
&lt;h2 id="write-prescriptive-prompts-for-complex-tasks"&gt;Write prescriptive prompts for complex tasks&lt;/h2&gt;&lt;p&gt;I created the personal productivity MCP server to help me take my favourite prompts from system to system. MCP (Model Context Protocol) servers provide a standardized way to expose tools and context to AI agents across different platforms. One thing I learned from my colleague Anand Murthy about how to write such prompts is to be extremely prescriptive about the actions and tools that I want the agent to use.&lt;/p&gt;
&lt;p&gt;Generic prompts like "help me debug this GitHub Actions workflow" leave too much room for interpretation. Instead, specify exact commands, tools, and steps. For example, if I'm looking to debug a GitHub Actions issue, the prompt that I have looks like this:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;You are helping me debug a failed GitHub Actions workflow. Follow these steps to systematically analyze and resolve the issue:

1. **Extract workflow information**: Parse the provided URL to identify:
   - Repository owner and name
   - Workflow run ID
   - Workflow name
   - Branch/commit that triggered the run

2. **Fetch workflow logs using GitHub CLI**:
   - Use `gh run list` to verify the workflow run exists
   - Use `gh run view &amp;lt;run-id&amp;gt;` to get detailed run information
   - Use `gh run view &amp;lt;run-id&amp;gt; --log` to download and display the full logs
   - Use `gh run view &amp;lt;run-id&amp;gt; --log-failed` to focus on failed job logs

3. **Analyze the failure**:
   - Identify which job(s) failed and at what step
   - Look for error messages, exit codes, and stack traces
   - Check for common issues: dependency problems, permission errors, timeout issues, resource constraints
   - Examine the workflow configuration and environment setup

4. **Provide debugging guidance**:
   - Explain what went wrong in simple terms
   - Suggest specific fixes or configuration changes
   - Provide commands or code snippets to resolve the issue
   - Recommend preventive measures to avoid similar failures

5. **Context-aware solutions**:
   - Consider the project type (Python, Node.js, etc.) and suggest appropriate fixes
   - Check for recent changes that might have caused the failure
   - Suggest workflow improvements or optimizations

6. **Follow-up actions**:
   - Recommend next steps for testing the fix
   - Suggest monitoring or alerting improvements
   - Provide guidance on preventing similar issues

Workflow URL: {workflow_url}

Focus on providing actionable, specific solutions rather than generic troubleshooting advice. Use the GitHub CLI commands to gather comprehensive information about the failure.
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Notice how prescriptive this prompt is. Rather than being a generic troubleshooting guide, it's a step-by-step guide that the agent can follow, down to the level of exact CLI commands to run. Critically, those CLI commands (&lt;code&gt;gh run list&lt;/code&gt;, &lt;code&gt;gh run view&lt;/code&gt;) are commands that I have auto-approved in my IDE, so the agent can execute the entire workflow autonomously without interrupting me for approval at each step.&lt;/p&gt;
&lt;p&gt;The prompt was written with AI assistance, which allows me to iterate to the level of detail I want with minimal effort. I start with a rough outline, then ask the agent to make it more specific, add command examples, and refine the steps until it's actionable enough for autonomous execution.&lt;/p&gt;
&lt;h2 id="use-plan-mode-for-complex-tasks"&gt;Use plan mode for complex tasks&lt;/h2&gt;&lt;p&gt;Plan mode in Cursor and Claude significantly improves agent performance on complex tasks. Users of AI-assisted coding tools consistently report that plan mode helps agents stay on course, compared to agents working without a structured plan. This mirrors how humans perform better with explicit plans.&lt;/p&gt;
&lt;p&gt;The mechanism is straightforward: the agent first generates a detailed plan, you review and refine it, then the agent executes against that plan. This separation of planning and execution prevents the agent from going down rabbit holes or making premature implementation decisions.&lt;/p&gt;
&lt;p&gt;In my experience, agents often complete tasks in one attempt after a few iterations on a well-defined plan. The key is ensuring the plan is specific and properly scoped before execution begins. I've found that plans work best when they include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Specific files and functions to modify&lt;/li&gt;
&lt;li&gt;Clear acceptance criteria&lt;/li&gt;
&lt;li&gt;Dependencies and ordering constraints&lt;/li&gt;
&lt;li&gt;Test cases or validation steps&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Without this structure, agents tend to make assumptions, skip steps, or get distracted by tangential improvements.&lt;/p&gt;
&lt;h2 id="managing-multiple-background-agents"&gt;Managing multiple background agents&lt;/h2&gt;&lt;p&gt;Multiple background agents can be powerful, but they require careful management. Unless agents are handling mundane, well-defined tasks, context switching between multiple active agents becomes challenging. At that point, you're operating at the speed of thought, which requires significant cognitive overhead.&lt;/p&gt;
&lt;p&gt;I've found that multiple agents work well when they're working on independent, well-scoped tasks. For example, one agent might be researching documentation while another refactors a specific module. But when tasks have dependencies or require coordination, a single agent with a clear plan tends to perform better than multiple agents trying to coordinate.&lt;/p&gt;
&lt;p&gt;The cognitive load turns out to be more than keeping track of what each agent is doing; we also need to ensure they don't conflict with each other. Two agents modifying the same file simultaneously, or one agent's changes breaking assumptions another agent made, creates more problems than it solves.&lt;/p&gt;
&lt;h2 id="additional-resources"&gt;Additional resources&lt;/h2&gt;&lt;p&gt;Others have written extensively about effective coding agent workflows. Here's a curated collection of resources I've found valuable:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Oct/25/coding-agent-tips/"&gt;Simon Willison's coding agent tips&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.geoffreylitt.com/2025/10/24/code-like-a-surgeon"&gt;Geoffrey Litt suggests coding like a surgeon&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://x.com/omarsar0/status/1984641893519839271"&gt;&lt;code&gt;@omarsar0&lt;/code&gt; on Twitter loves plan mode on Claude Code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://x.com/mattpocockuk"&gt;&lt;code&gt;@mattpocockuk&lt;/code&gt; has awesome tips on how to use AI for coding&lt;/a&gt;, including &lt;a href="https://x.com/mattpocockuk/status/1985056806893211915"&gt;this tip&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Nov/6/async-code-research/"&gt;Simon Willison (again!) on async code research with coding agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.linkedin.com/posts/sebastian-wallkoetter_my-favourite-question-to-spot-a-vibe-coder-activity-7394726959349592064--baU"&gt;Sebastian Wallkötter on preventing AI spaghetti through intermediate reviews&lt;/a&gt;: The key insight is that AI coding's bottleneck is code review, not code generation. Small mistakes compound when AI re-ingests its own errors as context. The solution: implement features in small increments with intermediate reviews, fixing "5 second issues" as you go, rather than letting mistakes accumulate into spaghetti code that takes hours to untangle.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What are your tips for safe ways to let your coding agent work autonomously? And what did you like most about this post? Let me know in the comments below!&lt;/p&gt;
</content></entry><entry><title>Use coding agents to write Marimo notebooks</title><link href="https://ericmjl.github.io/blog/2025/10/28/use-coding-agents-to-write-marimo-notebooks/" rel="alternate"/><updated>2025-10-28T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:820f6ef9-cdc4-384d-a08e-890efd1130b9</id><content type="html">&lt;p&gt;If you're like me, you might find coding with AI assistants somewhat addictive. And if you're like me, you might also like to write code in Marimo notebooks, the modern alternative to Jupyter that offers better reproducibility and cleaner Python development.&lt;/p&gt;
&lt;p&gt;Turns out there's a way to put these two together for automated Python development and data science workflows, creating a powerful combination for rapid prototyping and iterative coding.&lt;/p&gt;
&lt;h2 id="marimo-s-watch-flag"&gt;Marimo's &lt;code&gt;--watch&lt;/code&gt; Flag&lt;/h2&gt;&lt;p&gt;A few months ago, at SciPy 2025, my friend &lt;a href="https://trevorma.nz/"&gt;Trevor Manz&lt;/a&gt; showed me a cool neat trick for writing Marimo notebooks. Apart from launching a Marimo notebook in &lt;a href="https://docs.marimo.io/guides/package_management/inlining_dependencies/"&gt;sandbox mode&lt;/a&gt;, you add a &lt;code&gt;--watch&lt;/code&gt; flag:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;uvx&lt;span class="w"&gt; &lt;/span&gt;marimo&lt;span class="w"&gt; &lt;/span&gt;edit&lt;span class="w"&gt; &lt;/span&gt;--sandbox&lt;span class="w"&gt; &lt;/span&gt;my_notebook.py&lt;span class="w"&gt; &lt;/span&gt;--watch
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;When edits are made to the source file &lt;code&gt;notebook.py&lt;/code&gt;, they will now be reflected in the browser as well. This was my reaction:&lt;/p&gt;
&lt;p&gt;&lt;img src="minion-what.webp" alt="minion-what.webp"&gt;&lt;/p&gt;
&lt;p&gt;If you ever meet Trevor in person, he can confirm that reaction of mine.&lt;/p&gt;
&lt;h2 id="ensure-code-quality-with-marimo-check"&gt;Ensure code quality with &lt;code&gt;marimo check&lt;/code&gt;&lt;/h2&gt;&lt;p&gt;So now, AI coding assistants can write your Marimo notebooks for you... but it's not always going to be correct first time, right? After all, the latest features of Marimo are not going to be part of the large language model training sets.&lt;/p&gt;
&lt;p&gt;Turns out, Marimo also ships with a &lt;code&gt;check&lt;/code&gt; command that you can ask coding agents to call on:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;uvx&lt;span class="w"&gt; &lt;/span&gt;marimo&lt;span class="w"&gt; &lt;/span&gt;check&lt;span class="w"&gt; &lt;/span&gt;my_notebook.py
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And that will print to stdout any issues that Marimo finds that break its execution model, such as variables that are repeated variables or invalid cells.&lt;/p&gt;
&lt;p&gt;You can instruct coding agents to always run &lt;code&gt;marimo check&lt;/code&gt; by adding the following prompt (or analogous) into &lt;code&gt;AGENTS.md&lt;/code&gt;:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;When editing Marimo notebooks, always run &lt;span class="sb"&gt;`uvx marimo check`&lt;/span&gt; on the file and fix all issues that you find.
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This will virtually guarantee correctly-written, AI-generated notebooks. All that's left for us as users is to check the correctness of the analysis that was done.&lt;/p&gt;
&lt;h2 id="real-world-use"&gt;Real-world use&lt;/h2&gt;&lt;p&gt;Now, AI coding assistants (like Cursor, GitHub Copilot, or Claude Code) can write and edit large chunks of Marimo notebook cells for you, check what they wrote, and fix any syntactic issues that show up. And by checking that the cells are syntactically valid. Now you can speed-run those routine and yet highly mundane data manipulation code-writing activities while making yourself an espresso drink. This aligns perfectly with my philosophy on &lt;a href="../../../../2019/3/20/how-i-work/"&gt;optimizing for productivity in data science workflows&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I've used this mode to speed-run first versions of &lt;a href="../../../4/3/bayesian-superiority-estimation-with-r2d2-priors-a-practical-guide-for-protein-screening/"&gt;probabilistic models in PyMC&lt;/a&gt;, create explainer notebooks for hard concepts, make notebooks that process data, and many, many more things that you'd usually be able to do within a coding notebook system. The key thing that makes this work is feedback given (via the command line) that the coding agent can use for self-correction.&lt;/p&gt;
&lt;h2 id="advanced-functionality-using-mcp-and-built-in-ai-features"&gt;Advanced functionality using MCP and built-in AI features&lt;/h2&gt;&lt;p&gt;It doesn't stop there, though. There's a new &lt;code&gt;--mcp&lt;/code&gt; flag that makes a notebook an MCP server that coding agents can connect to; read more about it &lt;a href="https://opensourcedev.substack.com/p/beyond-chatbots-how-i-turned-python"&gt;here&lt;/a&gt;. Marimo also has built-in AI editing capabilities itself as well. Check out the functionality &lt;a href="https://docs.marimo.io/guides/editor_features/ai_completion/#custom-copilots"&gt;here&lt;/a&gt;, as well as Vincent Warmerdam's short video on &lt;a href="https://www.youtube.com/shorts/CnHOGE46x3o"&gt;using coding agents from &lt;em&gt;within&lt;/em&gt; Marimo&lt;/a&gt;. He's got my vote for best facial/eyebrow expressions from a coding YouTuber!&lt;/p&gt;
&lt;h2 id="addendum"&gt;Addendum&lt;/h2&gt;&lt;p&gt;After sharing this post on LinkedIn, Séverin H. shared a couple of additional use cases worth highlighting:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;One use case I would also recommend is getting the coding assistant to run queries for you, especially when it is to debug a existing query. You can ask [it] to check for corner cases (and especially dig into the data to understand the corner cases).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(&lt;a href="https://www.linkedin.com/feed/update/urn:li:activity:7391852642743775232?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A7391852642743775232%2C7391858466161606656%29&amp;amp;dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287391858466161606656%2Curn%3Ali%3Aactivity%3A7391852642743775232%29"&gt;LinkedIn comment&lt;/a&gt;)&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;The &lt;code&gt;--watch&lt;/code&gt; flag is indeed very interesting use case. Also to note they created a Claude.md to get you started that you can directly curl:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;curl&lt;span class="w"&gt; &lt;/span&gt;https://docs.marimo.io/CLAUDE.md&lt;span class="w"&gt; &lt;/span&gt;&amp;gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;your&lt;span class="w"&gt; &lt;/span&gt;agents.md&lt;span class="w"&gt; &lt;/span&gt;file&lt;span class="o"&gt;]&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Some more reference from their blog: &lt;a href="https://marimo.io/blog/claude-code"&gt;https://marimo.io/blog/claude-code&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(&lt;a href="https://www.linkedin.com/feed/update/urn:li:activity:7391852642743775232?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A7391852642743775232%2C7391854914886533121%29&amp;amp;dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287391854914886533121%2Curn%3Ali%3Aactivity%3A7391852642743775232%29"&gt;LinkedIn comment&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;Thanks for the suggestions, Séverin!&lt;/p&gt;
</content></entry><entry><title>Exploring Skills vs MCP Servers</title><link href="https://ericmjl.github.io/blog/2025/10/20/exploring-skills-vs-mcp-servers/" rel="alternate"/><updated>2025-10-20T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:cb890a27-6c6a-351a-b47a-c3db04a3f25d</id><content type="html">&lt;p&gt;I spent time digging through Anthropic's skills repository. These are my first impressions, organized for clarity and future reference.&lt;/p&gt;
&lt;h2 id="what-the-anthropic-skills-repository-offers"&gt;What the Anthropic Skills repository offers&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Creative &amp;amp; design workflows&lt;/strong&gt;: &lt;code&gt;algorithmic-art&lt;/code&gt; (generative art with p5.js), &lt;code&gt;canvas-design&lt;/code&gt; (beautiful PNG/PDF outputs guided by design philosophies), &lt;code&gt;theme-factory&lt;/code&gt; (pre-set or on-the-fly themes), and &lt;code&gt;slack-gif-creator&lt;/code&gt; (animated GIFs tuned for Slack). These are turnkey “taste plus tooling” bundles that let the model produce high-quality visuals with consistent aesthetics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Document skills for real formats&lt;/strong&gt;: &lt;code&gt;document-skills/&lt;/code&gt; cover &lt;code&gt;pptx&lt;/code&gt;, &lt;code&gt;docx&lt;/code&gt;, &lt;code&gt;pdf&lt;/code&gt;, and &lt;code&gt;xlsx&lt;/code&gt; with serious capabilities: layout/templates, tracked changes and comments, text/table extraction, merges/splits, charting, formulas, and formatting preservation. This feels like a pragmatic spec+runtime for working with binary formats—lean instructions up front, heavy lifting when needed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Development &amp;amp; technical utilities&lt;/strong&gt;: &lt;code&gt;artifacts-builder&lt;/code&gt; (compose complex Claude HTML artifacts using React/Tailwind/shadcn), &lt;code&gt;webapp-testing&lt;/code&gt; (Playwright-driven UI testing), and &lt;code&gt;mcp-builder&lt;/code&gt; (guidance for creating high-quality MCP servers). These reduce boilerplate for the “build and test” loop.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Enterprise &amp;amp; communication&lt;/strong&gt;: &lt;code&gt;brand-guidelines&lt;/code&gt; (apply Anthropic’s official brand colors and typography) and &lt;code&gt;internal-comms&lt;/code&gt; (status reports, newsletters, FAQs). These encode editorial and brand guardrails so outputs stay on-message.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Meta skills and templates&lt;/strong&gt;: &lt;code&gt;skill-creator&lt;/code&gt; and &lt;code&gt;template-skill&lt;/code&gt; show how to structure your own skills: a folder per skill with a &lt;code&gt;SKILL.md&lt;/code&gt; (YAML front matter for &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt;, plus instructions/examples/guidelines), optional scripts, and assets. This is the pattern to replicate.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you want the source for these examples, it’s viewable in the repo. Start here: &lt;code&gt;https://github.com/anthropics/skills&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="how-skills-are-loaded-and-used"&gt;How skills are loaded and used&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Minimal prompt footprint&lt;/strong&gt;: A skill's short description is passed up front. The larger &lt;code&gt;skill.md&lt;/code&gt; is only read when the model decides it needs more detail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;On-demand details&lt;/strong&gt;: The model can iterate (ReAct loop) to fetch instructions and then execute scripts or read additional files.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This access pattern keeps the initial token budget small and defers detail until it’s actually needed.&lt;/p&gt;
&lt;h2 id="contrast-with-mcp-servers"&gt;Contrast with MCP servers&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;MCP call shape&lt;/strong&gt;: Tool names and descriptions are typically sent on every call. That keeps tools globally discoverable but increases token overhead.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Skills call shape&lt;/strong&gt;: A tiny descriptor up front; details fetched lazily. Lower baseline token cost.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Distribution model&lt;/strong&gt;:&lt;ul&gt;
&lt;li&gt;MCP: Centrally hostable (e.g. web server) or vendable (e.g., a Python package). Easy to version, release, and update for many users at once.&lt;/li&gt;
&lt;li&gt;Skills: Feel local-first. You can drag-and-drop into a Claude workspace. Easy to customize, but harder to standardize and propagate updates across a team.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Given current industry patterns, MCP servers are the widely accepted way to expose functionality to LLMs across tools and vendors. Skills are Anthropic-specific at the moment.&lt;/p&gt;
&lt;h2 id="token-efficiency-and-why-its-emphasized"&gt;Token efficiency (and why it’s emphasized)&lt;/h2&gt;&lt;p&gt;Anthropic’s materials lean into token efficiency. The cost of LLM calls adds up, and repeatedly sending long tool descriptions can be expensive. Skills reduce baseline tokens: spend a handful of tokens to register intent, read detail only when needed, then execute. That’s the economic story.&lt;/p&gt;
&lt;h2 id="practical-trade-offs"&gt;Practical trade-offs&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Standardization vs customization&lt;/strong&gt;:&lt;ul&gt;
&lt;li&gt;MCP servers: Strong for shared, versioned, and centrally updated capabilities.&lt;/li&gt;
&lt;li&gt;Skills: Great for rapid, local customization without infrastructure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Discovery vs cost&lt;/strong&gt;:&lt;ul&gt;
&lt;li&gt;MCP: High discoverability; the model always sees the tools. Higher token floor.&lt;/li&gt;
&lt;li&gt;Skills: Low token floor; details fetched when needed. Requires the model to choose to read more.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="open-questions-im-tracking"&gt;Open questions I’m tracking&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;How will teams distribute and update skills at scale without a central registry or packaging story?&lt;/li&gt;
&lt;li&gt;Will skills gain cross-vendor support, or remain Anthropic-only?&lt;/li&gt;
&lt;li&gt;What’s the best practice to map a complex skill into smaller, composable units without losing clarity?&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="early-take"&gt;Early take&lt;/h2&gt;&lt;p&gt;IMO, skills are a clear attempt to lower token costs and streamline task-specific workflows with minimal upfront context. MCP servers remain the well-understood, cross-ecosystem pattern for exposing capabilities. If your goal is a shareable, versioned interface for many users, MCP is still the safer default. If you need quick, local customization inside Claude with a lean prompt footprint, skills are compelling. But this field has been evolving at breawkneck speed anyways, so expect changes.&lt;/p&gt;
</content></entry><entry><title>How to expose any documentation to any LLM agent</title><link href="https://ericmjl.github.io/blog/2025/10/19/how-to-expose-any-documentation-to-any-llm-agent/" rel="alternate"/><updated>2025-10-19T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:70188213-fcc1-328b-8c1f-f1a5fa7f43e1</id><content type="html">&lt;p&gt;Like cars that lose value as soon as they roll off the lot, LLMs become outdated as soon as their training sets are fixed. Software documentation evolves constantly—new features, API changes, bug fixes, and best practices emerge daily. Yet AI agents are stuck with whatever knowledge was captured in their training data, creating a fundamental mismatch between what they know and what developers actually need in real-time.&lt;/p&gt;
&lt;p&gt;Building LlamaBot taught me something unexpected: the hardest part of AI-assisted development isn't writing better prompts or designing cleaner abstractions. It's equipping AI agents with up-to-date information in a stable, standardized fashion.&lt;/p&gt;
&lt;p&gt;Most developers know the frustration of context-switching between code and documentation. You're deep in a coding session, need to check how a specific function works, and suddenly you're hunting through static documentation files. AI agents face this same problem, but with an added layer of complexity—they need structured, queryable access to documentation that can be searched semantically.&lt;/p&gt;
&lt;p&gt;I discovered that web searches by coding agents were less reliable than manually adding context, but manual approaches don't scale. The solution emerged through the Model Context Protocol (MCP), a standard that enables LLMs to interact with external tools and data sources. In LlamaBot v0.13.10, I introduced a documentation MCP server that automatically equips AI agents with current information. This enables AI agents to access organizational knowledge, process documentation, and domain expertise in structured ways.&lt;/p&gt;
&lt;h2 id="the-obsolescence-problem-in-ai-assisted-development"&gt;The obsolescence problem in AI-assisted development&lt;/h2&gt;&lt;p&gt;The core issue more than mere documentation access, it's about obsolescence. LLMs are trained on data that becomes outdated the moment it's fixed in their training sets. Meanwhile, software documentation evolves constantly. New features are added, APIs change, bugs are fixed, and best practices emerge. Yet AI agents remain frozen in time, working with knowledge that may be months or years out of date.&lt;/p&gt;
&lt;p&gt;Consider a typical data science workflow: you're building an AI pipeline and need to understand how LlamaBot's StructuredBot handles data validation. The AI agent might reference documentation from six months ago, missing critical updates or new features that could solve your problem more elegantly. This creates a fundamental mismatch between what the agent knows and what's actually available.&lt;/p&gt;
&lt;p&gt;The deeper problem is that AI agents need structured, queryable access to documentation that can be searched semantically and updated automatically. They need to understand not just what functions exist, but how they relate to each other, what patterns they follow, and how they fit into broader workflows. Static documentation simply cannot provide this level of contextual understanding, particularly in data science environments where teams maintain scattered knowledge across wikis, Slack threads, and onboarding documents.&lt;/p&gt;
&lt;h2 id="building-a-semantic-documentation-layer"&gt;Building a semantic documentation layer&lt;/h2&gt;&lt;p&gt;LlamaBot's MCP server demonstrates how to give AI agents structured access to its documentation by creating a dynamic, queryable knowledge base that agents can search semantically. The &lt;a href="https://github.com/ericmjl/llamabot/blob/main/llamabot/mcp_server.py"&gt;implementation&lt;/a&gt; centers around a single tool:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nd"&gt;@mcp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;docs_search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Search through LlamaBot documentation and source code.&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;docstore&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_results&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;query&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;results&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This interface sits in front of a data pipeline that builds a vector database for the documentation. The server fetches the latest documentation from GitHub, extracts Python module docstrings from source code, and constructs a LanceDB vector database optimized for semantic search. The database is built during CI/CD and packaged directly with the wheel distribution, giving users instant access without setup while staying current with each release.&lt;/p&gt;
&lt;p&gt;This approach works with any AI agent system through the MCP protocol, providing a standardized way to keep AI agents current with documentation.&lt;/p&gt;
&lt;h2 id="the-architecture-behind-semantic-documentation"&gt;The architecture behind semantic documentation&lt;/h2&gt;&lt;p&gt;The MCP server combines several technologies to create a robust documentation system. FastMCP handles the protocol implementation, enabling seamless communication between AI agents and the documentation database. LanceDB powers the semantic search capabilities, leveraging LlamaBot's existing &lt;code&gt;LanceDBDocStore&lt;/code&gt; class with hybrid search and reranking for optimal results.&lt;/p&gt;
&lt;p&gt;The system uses the checked-out documentation from the repository during the CI/CD build process, ensuring the packaged database contains current information. The build script first attempts to fetch docs from GitHub, but falls back to the local &lt;code&gt;docs/&lt;/code&gt; directory when available, making it work seamlessly in both CI/CD and development environments. The build process runs the &lt;code&gt;scripts/build_mcp_docs.py&lt;/code&gt; script during CI/CD, which creates the LanceDB database and copies it to &lt;code&gt;llamabot/data/mcp_docs/&lt;/code&gt; for packaging.&lt;/p&gt;
&lt;p&gt;I believe this architecture represents a fundamental shift in how we think about documentation for AI systems. Instead of treating documentation as static reference material, we're creating dynamic, queryable knowledge bases that AI agents can interact with directly.&lt;/p&gt;
&lt;h2 id="the-core-pattern-to-replicate"&gt;The core pattern to replicate&lt;/h2&gt;&lt;p&gt;The LlamaBot MCP server follows a straightforward pattern that any package or documentation source can replicate. Here's the essential blueprint:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Build a semantic database during CI/CD&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Extract documentation from your source (GitHub, local docs, API references)&lt;/li&gt;
&lt;li&gt;Parse and chunk the content appropriately for your domain&lt;/li&gt;
&lt;li&gt;Create a vector database (LanceDB, Chroma, or similar) with semantic search capabilities&lt;/li&gt;
&lt;li&gt;Package the database with your distribution&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;2. Create an MCP server with a search tool&lt;/strong&gt;&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="nd"&gt;@mcp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nf"&gt;docs_search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Search through your documentation and source code.&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;docstore&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_results&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;query&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;results&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;3. Make it discoverable and configurable&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Provide a simple launch command (like &lt;code&gt;yourpackage mcp launch&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Include clear setup instructions for MCP-compatible tools&lt;/li&gt;
&lt;li&gt;Ensure the database updates automatically with each release&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;4. Design for your specific knowledge domain&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Include not just API docs, but process documentation, examples, and institutional knowledge&lt;/li&gt;
&lt;li&gt;Structure the content for semantic search rather than keyword matching&lt;/li&gt;
&lt;li&gt;Consider what context your users need most when working with AI agents&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="seamless-integration-with-modern-development-tools"&gt;Seamless integration with modern development tools&lt;/h2&gt;&lt;p&gt;The MCP server works with any MCP-compatible coding environment, including Cursor, VSCode, and other modern development tools. Configuration requires a single command in your MCP settings:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;uvx&lt;span class="w"&gt; &lt;/span&gt;--with&lt;span class="w"&gt; &lt;/span&gt;llamabot&lt;span class="o"&gt;[&lt;/span&gt;all&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;llamabot&lt;span class="w"&gt; &lt;/span&gt;mcp&lt;span class="w"&gt; &lt;/span&gt;launch
&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Once configured, AI agents can query LlamaBot documentation using natural language queries. Ask "How do I use StructuredBot for data extraction?" and the agent receives structured results with content, relevance scores, and metadata. This contextual information enables agents to provide accurate, up-to-date assistance without manual documentation lookup.&lt;/p&gt;
&lt;p&gt;This reduces context-switching between code and documentation. AI agents can access relevant information and provide suggestions based on current code patterns and usage examples. This approach is particularly valuable for data science teams who need to maintain consistency across experiments while leveraging the latest library capabilities.&lt;/p&gt;
&lt;h2 id="comparing-documentation-approaches-for-ai-agents"&gt;Comparing documentation approaches for AI agents&lt;/h2&gt;&lt;p&gt;There are several ways to provide documentation to AI agents, each with distinct trade-offs. Understanding these approaches helps clarify why the MCP server approach represents a significant improvement.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Native tool documentation&lt;/strong&gt; (like Cursor's built-in docs capabilities) offers seamless integration and can fetch docs from online sources, but you're limited to how the tool fetches those docs. It may not be able to access certain systems gated behind access controls or include custom organizational knowledge and process documentation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Manual repository inclusion&lt;/strong&gt; works well for users familiar with IDEs, workspaces, and development concepts, but it requires familiarity with these practices that are unfamiliar to non-technical users or individual developers. It also doesn't scale beyond individual developers. The documentation becomes part of the context window, consuming valuable tokens and potentially overwhelming the agent with irrelevant information.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Copy-paste or file upload&lt;/strong&gt; (like Claude Projects) provides flexibility for non-technical users but creates maintenance overhead. You must manually update documentation when it changes, and there's no semantic search capability—agents can only work with what you explicitly provide.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Web search by agents&lt;/strong&gt; seems convenient but creates inefficiency in the development workflow. If the documentation is up-to-date, the LLM will find it eventually, but it requires multiple iterations of web searches to locate the right information. I discovered this firsthand when building LlamaBot—web searches by coding agents required more iterations than directly providing context, but manual approaches don't scale.&lt;/p&gt;
&lt;p&gt;The MCP server approach provides automatic updates, semantic search, system-agnostic compatibility, and organizational knowledge integration. It offers a standardized way to keep AI agents current with evolving documentation. The trade-off is initial setup complexity, but this is mitigated by the pre-built databases that ship with packages.&lt;/p&gt;
&lt;h2 id="beyond-software-documentation-surfacing-any-process-knowledge"&gt;Beyond software documentation - surfacing any process knowledge&lt;/h2&gt;&lt;p&gt;The MCP approach extends beyond software documentation. The LlamaBot server gives an example of how organizations can surface their process documentation, institutional knowledge, and domain expertise.&lt;/p&gt;
&lt;p&gt;I believe that data science teams could transform their workflow documentation—experimental protocols, data validation procedures, model evaluation criteria, or deployment checklists—from scattered wikis, Slack threads, and buried onboarding documents into structured, queryable knowledge bases that AI agents can access and reference during development.&lt;/p&gt;
&lt;p&gt;I can see this approach scaling beyond individual libraries to entire organizational knowledge. We have all imagined AI agents that can query your team's coding standards, understand your deployment procedures, or reference your data governance policies—all without leaving their development environment. Each organization would maintain their own specialized knowledge base, creating networks of interconnected AI-accessible process documentation. How would one implement this? A documentation MCP server may be a great way to start.&lt;/p&gt;
&lt;p&gt;This approach isn't just for software docs. Imagine surfacing your team's process knowledge, onboarding guides, or even those golden nuggets buried in Slack threads. The MCP server pattern can turn scattered, informal knowledge into a living, searchable resource for both humans and AI agents, especially if you treat your processes as versioned software to be exposed to AI agents!&lt;/p&gt;
&lt;p&gt;In my experience, the most valuable knowledge in organizations often exists in informal channels—Slack conversations, email threads, or tribal knowledge that never gets documented. I believe the MCP approach provides a framework for capturing and surfacing this knowledge in ways that AI agents can understand and reference.&lt;/p&gt;
&lt;h2 id="the-future-of-ai-assisted-development"&gt;The future of AI-assisted development&lt;/h2&gt;&lt;p&gt;Future iterations could include real-time updates that rebuild databases when documentation changes, cross-organizational knowledge graphs, and usage pattern analysis that learns from how teams implement processes.&lt;/p&gt;
&lt;p&gt;The goal is to make AI agents active participants in organizational processes, capable of understanding team workflows and providing context-aware recommendations.&lt;/p&gt;
&lt;p&gt;This vision requires rethinking how we structure and maintain organizational knowledge. Instead of writing documentation solely for human consumption, we need to design knowledge systems that serve both human team members and AI agents, creating a symbiotic relationship between human creativity and AI capability while preserving institutional knowledge in accessible, queryable formats.&lt;/p&gt;
&lt;h2 id="getting-started-with-semantic-documentation"&gt;Getting started with semantic documentation&lt;/h2&gt;&lt;p&gt;The MCP server is available in LlamaBot v0.13.10 and later. Getting started requires minimal setup:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Install LlamaBot with MCP support: &lt;code&gt;pip install llamabot[all]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Configure your coding tool to use the MCP server&lt;/li&gt;
&lt;li&gt;Begin coding with AI agents that understand LlamaBot's capabilities&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The documentation database ships pre-built with the package, eliminating setup friction. The server exposes a single &lt;code&gt;docs_search&lt;/code&gt; tool that agents can use to find relevant documentation and source code information, creating a seamless development experience.&lt;/p&gt;
&lt;p&gt;This approach makes documentation an integral part of the AI agent's toolkit, resulting in more capable assistants that can help developers work more effectively.&lt;/p&gt;
&lt;p&gt;The future of AI-assisted development involves better integration between AI agents and the tools they need. LlamaBot's MCP server demonstrates how this integration can work in practice.&lt;/p&gt;
</content></entry><entry><title>A practical comparison of DSPy and LlamaBot for structured LLM applications</title><link href="https://ericmjl.github.io/blog/2025/10/18/a-practical-comparison-of-dspy-and-llamabot-for-structured-llm-applications/" rel="alternate"/><updated>2025-10-18T00:00:00Z</updated><author><name>Eric J. Ma</name></author><id>urn:uuid:0282121b-3e33-3ecb-9837-afd6b1121706</id><content type="html">&lt;p&gt;When Omar Khattabe presented &lt;a href="https://dspy.ai"&gt;DSPy 3.0&lt;/a&gt; at PyData Boston Cambridge last week, I finally had the chance to dig into a framework that's been generating significant buzz in the LLM development community. As someone who's built structured LLM applications with &lt;a href="https://ericmjl.github.io/llamabot/"&gt;LlamaBot&lt;/a&gt;, I was particularly curious about DSPy's core claim: that signatures represent the only abstraction you need for LLM-powered programs.&lt;/p&gt;
&lt;p&gt;The presentation focused on two key concepts: signatures as a new LLM abstraction and prompt optimization techniques. But what caught my attention was the practical similarity between DSPy's approach and what I've been doing with LlamaBot's StructuredBot. This led me to build a direct comparison using a real-world example from my personal expense tracking application.&lt;/p&gt;
&lt;h2 id="the-structured-llm-challenge"&gt;The structured LLM challenge&lt;/h2&gt;&lt;p&gt;Most developers working with LLMs face the same fundamental problem: how do you reliably extract structured data from unstructured inputs? Whether you're processing receipts, parsing documents, or analyzing text, you need consistent, typed outputs that integrate cleanly with your existing systems.&lt;/p&gt;
&lt;p&gt;Traditional approaches rely heavily on natural language prompts, which are fragile, hard to maintain, and difficult to optimize. DSPy proposes a different path through its signature abstraction, claiming this eliminates the need for verbose prompt engineering.&lt;/p&gt;
&lt;h2 id="a-real-world-comparison-receipt-processing"&gt;A real-world comparison: Receipt processing&lt;/h2&gt;&lt;p&gt;To test DSPy's claims, I built a practical comparison using an expense extraction system I developed for personal use. This application processes receipts in various formats (PNG, PDF, JPG, WEBP) and automatically extracts structured expense data into Notion — essentially a lightweight alternative to enterprise expense management systems.&lt;/p&gt;
&lt;p&gt;The challenge here is typical of structured LLM applications: converting unstructured visual and textual data into consistent, typed outputs that integrate with existing workflows. Let's see how both frameworks handle this task.&lt;/p&gt;
&lt;h3 id="llamabot-s-structuredbot-approach"&gt;LlamaBot's StructuredBot approach&lt;/h3&gt;&lt;p&gt;LlamaBot uses Pydantic models to define structured outputs, leveraging Python's type system for validation and documentation. The approach emphasizes explicit data modeling with detailed field descriptions:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;pydantic&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;enum&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;typing&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;pathlib&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;llamabot&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;as&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;lmb&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;FlowType&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;MONEY_OUT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Money Out&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;MONEY_IN&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Money In&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;TypeEnum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;PAYMENT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Payment&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;INVOICE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Invoice&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;PaymentMethodEnum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;CASH&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Cash&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;BANK_TRANSFER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Bank Transfer&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;CREDIT_CARD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Credit Card&amp;quot;&lt;/span&gt;
    &lt;span class="n"&gt;CHECK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;Check&amp;quot;&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;ExpenseData&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;transaction_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Short, memorable description of the purchase. E.g.: &amp;#39;Anker Dock&amp;#39;, &amp;#39;Coffee at Triangle Bar&amp;#39;, &amp;#39;dbrand laptop skin&amp;#39;&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;date&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;transaction date&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;transaction amount&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;category&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Business category, e.g. Office Supplies, Travel, Meals&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;TypeEnum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Either Payment or Invoice&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;flow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;FlowType&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Either &amp;#39;Money Out&amp;#39; or &amp;#39;Money In&amp;#39;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;payment_method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;PaymentMethodEnum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;How the payment was made.&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Brief business purpose or description of the expense.&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;reference_number&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Invoice/receipt number if visible&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;person&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="s2"&gt;&amp;quot;Person responsible or who made the purchase if mentioned.&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Usage&lt;/span&gt;
&lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lmb&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StructuredBot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;pydantic_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ExpenseData&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;model_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ollama_chat/gemma3n:latest&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;/path/to/receipt.png&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;h3 id="dspy-s-signature-approach"&gt;DSPy's signature approach&lt;/h3&gt;&lt;p&gt;DSPy takes a different approach with its signature abstraction, which defines both inputs and outputs in a single class. The framework emphasizes simplicity and automatic prompt optimization:&lt;/p&gt;
&lt;div class="hll"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;dspy&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nc"&gt;ExpenseExtraction&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Signature&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;Extract expense information from receipt images.&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;

    &lt;span class="n"&gt;receipt_image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Image&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;InputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Receipt image&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;transaction_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Short description of the purchase&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Transaction date (YYYY-MM-DD)&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Total transaction amount (number, no currency symbols)&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Business category (e.g., Office Supplies, Travel, Meals)&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nb"&gt;type&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Transaction type, either &amp;#39;Payment&amp;#39; or &amp;#39;Invoice&amp;#39;&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;flow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Cash flow direction, either &amp;#39;Money Out&amp;#39; or &amp;#39;Money In&amp;#39;&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;payment_method&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;How the payment was made (e.g., Cash, Bank Transfer, Credit Card, Check)&amp;quot;&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;purpose&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Brief business purpose or description&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;reference_number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Invoice/receipt number if present&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;default&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;None&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;person&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;OutputField&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;desc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Person involved, if mentioned&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;default&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;None&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Usage&lt;/span&gt;
&lt;span class="n"&gt;lm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;LM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;ollama_chat/gemma3n:latest&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;configure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;module&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;dspy&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ExpenseExtraction&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;module&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;receipt_image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;images&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="comparing-the-approaches"&gt;Comparing the approaches&lt;/h2&gt;&lt;p&gt;Both frameworks successfully extracted structured data from receipt images, but they take fundamentally different approaches to the problem.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;LlamaBot's StructuredBot&lt;/strong&gt; leverages Python's existing type system through Pydantic models. This approach provides several advantages: automatic validation, IDE support, and integration with existing Python data processing pipelines. The explicit type definitions make the data contract clear and enforceable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DSPy's signatures&lt;/strong&gt; offer a more streamlined interface that combines input and output definitions in a single class. The framework's strength lies in its automatic prompt optimization capabilities, which can improve performance over time without manual intervention.&lt;/p&gt;
&lt;h2 id="key-differences-in-practice"&gt;Key differences in practice&lt;/h2&gt;&lt;p&gt;The most noticeable difference is verbosity. LlamaBot requires more explicit type definitions and imports, while DSPy's signature approach is more concise. However, this conciseness may come at the cost of some type safety and IDE support that Pydantic provides.&lt;/p&gt;
&lt;p&gt;Both frameworks use LiteLLM for model routing, making it easy to switch between different LLM providers. The model configuration syntax is identical, which suggests a common underlying architecture.&lt;/p&gt;
&lt;h2 id="the-schema-first-principle"&gt;The schema-first principle&lt;/h2&gt;&lt;p&gt;Regardless of which framework you choose, structured LLM applications require careful upfront schema design. The bulk of development time goes into defining your data model, not writing prompts. This schema-first approach is what makes these frameworks powerful—they force you to think clearly about your data requirements before implementation.&lt;/p&gt;
&lt;h2 id="looking-ahead-dspy-s-broader-vision"&gt;Looking ahead: DSPy's broader vision&lt;/h2&gt;&lt;p&gt;DSPy's claim that signatures are the only abstraction needed for LLM applications is ambitious but not entirely accurate. The framework includes additional abstractions like modules and optimizers that handle more complex scenarios. Signatures represent the core abstraction for simple input-output transformations, but building production LLM applications often requires more sophisticated orchestration.&lt;/p&gt;
&lt;p&gt;I'm planning to explore DSPy's more advanced features as I rebuild LlamaBot's agent abstractions. The goal is to understand how to construct autonomous LLM agent frameworks rather than individual agents—a challenge that requires thinking beyond simple input-output mappings.&lt;/p&gt;
&lt;p&gt;Being unfamiliar with DSPy's documentation initially, I found it challenging to follow, but thanks to fellow PyData Boston Cambridge organizer &lt;a href="https://www.linkedin.com/in/nnssa/"&gt;Nash Sabti&lt;/a&gt;'s guidance, I was able to make it happen and build this comparison.&lt;/p&gt;
&lt;p&gt;The structured LLM landscape is rapidly evolving, and frameworks like DSPy and LlamaBot are pushing the boundaries of what's possible. The key insight is that successful LLM applications require the same engineering discipline as traditional software: clear interfaces, robust error handling, and maintainable abstractions.&lt;/p&gt;
</content></entry></feed>