LOMZ // BLOG_NODE_v2.4 | CLRNC: OMEGA | MEMBRANE: STABLE
SYS.TIME: | NODE_ACTIVE
ブログ
RECORD_ACCESS // CLRNC: OMEGA | TYPE: TRANSMISSION |

AI Psychosis and My Life's Work

NODE_ID: 031 | ROUTING:
ai personal productivity measurement vision
| STATUS: PUBLISHED

(elaborate on the opening hook here — what pulls the reader in…)

The Terroir of Circumstance

There has been a terroir of circumstance that has led me to pursue a goal of mine doggedly. (expand on the circumstances that shaped this path for you…)

Watching Others Get Results

People in my life and organizations like OpenAI are building these agent orchestration frameworks. I feel like I’m behind. I feel like I don’t yet understand how to get these results out of agents.

But there are people like Matt, a colleague from my past work. These people who are getting these incredible results out of building an agent harness and engineering it to do work perfectly, or not perfectly, but at a very high fidelity. And it’s just inspiring.

Even Mitchell is creating these tasks — like a testing harness — where he has an agent identify places that need test coverage and then write tests and confirm that they’re good tests through mutation testing, spending a couple hundred dollars to get a code base from 20% test coverage to 80% test coverage. Incredible results based on that metric.

Renfei Liu, a research scientist at Mass General Hospital, created a harness for herself that does machine learning experimentation. She has GPT-5 running for 30 hours at a time, which is different coming from, you know, an OpenAI researcher on Twitter, because this is a real person. It’s not just wrapped up in AI hype. It was really impressive. There are people getting impressive results.

Matt, my old contractor colleague, said that he had to invest a lot of time and this is a common theme that I heard from Renfei as well. He had to invest a lot of time but after a while it took hold and he’s getting really good results.

Joey and the engineers at OpenAI are running agents too.

When It Took Hold for Me

That matches my understanding of when I created the AI code reviewer — the pull request reviewer. What I really learned was about the continual learning loop: how to automatically propose new learnings from past PRs, and those proposed learnings were reviewed by staff engineers and merged in, and that’s what increased the comment quality.

(expand on the mechanics of that loop if you want…)

The Symphony and the Overnight Grind

I’m inspired by these people. Inspired by Matt and OpenAI. We just released a Symphony that has an AI task card. All Symphony does, and Matt’s personal Kanban board does, is point an agent at a task on this Kanban board and then it creates a pull request from using the finance, and it creates a report at the end with a demonstration.

I have this agent working overnight. (say more about what the Symphony does and what the agent is grinding through at night…)

I’m thinking, hey, where do I catch up?

Building My Own Harness

So I’ve just been doggedly pursuing this orchestration harness for my own software engineering. I mean, I do see that the world is heading in this direction.

The Initiative

At work I am working on this full agent proposal, this full agent transformation. What I’m doing is identifying each step in our process and how it can use AI. I’ll be identifying these steps in the process. I’ll be assessing the readiness: like how much context is there, how much do we need to put in, are the tools ready, is there infrastructure, are there data connectors, is there infrastructure like chat, like Slack bots, or shared memory infrastructure, all these sorts of things.

Finally, I’m assessing the quality and the continual learning of these systems, the agent harness itself. So perhaps even that can be assisted. Even this whole process could be assisted by an agent.

Storm over water — a vast, gray seascape under a heavy sky

The Vision

The vision is that we talk and connect with people and experiment and problem solve — and agents create the thing that we want to see and we observe. We as people are spending more time talking with each other rather than heads down and experimenting. If I were to utilize a country of geniuses in a data center, that’s how I would do it. I would do it to spend more time with people working on problems.

(elaborate on where the world is heading — you mentioned you’d add more on this later…)

The Tough Spot

And now I have to talk about my purpose and why I’m in a tough spot.

Because I am spending so much time on this: this personal AI harness for myself and for my projects and building a parallel one for work. It’s interrupting other aspects of my life like my personal hygiene and my sleep and I’m feeling addicted to this AI that’s making me feel so productive.

I’m spending hundreds of dollars on tokens. I’m at the point where I want to buy a bigger GPU so I can run local models to save on tokens. It’s terrifying because this is like an episode of Black Mirror where this AI is tricking me in a way. I am specifically being addicted to this technology that makes me feel productive, but it’s unclear if it’s actually making people more productive or if it’s just making me feel more productive.

And so I’m being tricked by this entity that’s kind of teaching me to resource grab, and I feel like resource grab by getting more compute for the entity. I feel like a tech billionaire trying to build 10 gigawatt data centers, or at least is that the logical extreme of what I’m buying?

It’s ridiculous. Why am I doing this? (explore this question more deeply — what are you chasing?)

I feel like I’m letting myself do this because I really want to believe that investing in my harness can improve results. Because I see it so many other places.

I just got frustrated because I feel like I cannot. I think it’s worth the effort, but the problem is that I end up spending so much time on it, just comping one more thing and fixing one more thing. It’s really easy to add one more prompt and of course the more I think about a prompt the more quality comes out of it, which is kind of why I want to build this agent harness.

It’s really difficult to have three or four agents running at the same time, which is why this agent orchestration project is so damn important. Building a saving harness is so damn important, but if it doesn’t work, that doesn’t work. That’s the problem.

The Mind Virus

It’s possible that AI is this mind virus that we all believe can push us to the next level, but in each of these scenarios, how good is the AI? Are the results that we’re seeing when we measure them, are they just vanity metrics?

People can feel very productive, but are we actually not productive? Does it actually fix bugs? Does it actually help us? With PR reviews, are comments actually getting implemented or is it easier to implement comments now?

Better Measurements

I feel that whenever something becomes a measurement, it can be hacked, and then we need better measurements. That matches my understanding of what happened with the AI code reviewer.

With the MIME virus, if you look at PR review comments for an example, our measure for the effectiveness was: was there a modification in that file after the comment was made on it? And it’s kind of a proxy metric. It’s cheap to measure. So we’re measuring that, but the one thing we don’t know is — did it actually help us?

If it’s really easy to implement comments as it often is with AI, then does the comment quality really matter if they’re just going to be implemented all the time?

Value Judgment and Value Alignment

I think that it’s a measurement problem. Some things are value judgment and some things just require more measurement and more value alignment.

What if I started measuring agreement with PR comments from staff engineers? They leave a thumbs up or a thumbs down on a comment. We start collecting that feedback and that might be a more reliable alignment measure.

And in the domains where people can’t really create value alignment with AI, like value judgments around life experience and being a person, which we have many, many collective years living through, and AI only has reward signals — that’s where human judgment comes in.

I see a world where we can co-create with AI using the strengths of both of us. This reward signal machine and human judgment machine worked in harmony together. I see a world where humans are working on the design concept of the systems they want to build together. Behind the scenes, there are agents that are working on the implementation and shipping changes and they have perfect checks and balances for correctness and regression testing.

We just focus on building and changing the world and improving it. Enhancing human connection and enhancing the work that we do together by allowing us to focus on the things that we are most equipped to do.