I think an underappreciated aspect is that we previously amortized the reading and comprehending over the time we spent designing and writing. I still find myself looking at changes expecting that I can review it in a few minutes when I am in effect coming in cold. That part of the job has changed, but fwiw it isn't so different than the experience I had reviewing another team's changes when I didn't know their services.
I have found the right harness helps, if you have a setup that gives you memory, specs and a graph-based understanding (that part is key) you can be very productive. For me, that's because being able to understand the big picture helps, the challenge of reading the code is always there but starting with 'this method/class/etc does X, so that means this is doing Y..." makes it easier and lets you come in warmer.
I generally produce nearly the same code I'd write myself about 5x faster with AI. I don't just let Claude Code run wild for a long time and have a mess to review. I have it do small chunks I can quickly review, give it feedback, iterate, etc until I like the output, then I move on to the next step. This takes time of course but I've found it faster than reviewing a giant mess of a code review.
My experience as well. Opus 5.5 is really good at matching the existing coding style. I often look at code and think I wrote it, only to realize it was written by Opus.
The trick is to give it enough guidance about the style you prefer.
We have multiple apps, but I'll share only the free ones; the others are B2B (they haul in major revenue for us, and businesses dealing with us will not like us disclosing this here), largest free one is MacroCodex with 17,000+ users.
We do not have this problem; we heavily use the models from OpenAI and Anthropic, but we have a custom harness built for ultimate cache efficiency and a purpose tuned workflow.
Codex is GREAT but it keeps updating and its UI changes often (remember mini git sidebar ui?)
Our Custom Harness does multiple things, like having a "Skill Selector," which uses our Gambler v1 (26B decision model responds in <100ms locally) and overlays tools on top of existing ones (as the models are RLVL over the specific tool in their harness, we overlay those tools with a familiar interface or format and inject custom features we want the model to interpret). Other than this, it leverages local models to accelerate development. Also, instead of providing custom tools, we actually use the same tools that their original model-provided harness uses. We've a list of 100s of such optimizations; these are just a few from the top of my head.
I have been working on harnesses for a very long time; I read research papers and implement them.
The apps use Flutter and Rust for all of our apps; no ads, no subscriptions.
So, macrocodex is basically a neural network product which "predicts" maintenance calories at a very high accuracy, https://macrocodex.app/
MacroCodex beats everything in its category and it does it for free, no ads, no subscriptions.
Symbiote and CalorieCodex will soon follow the suite.
If it fails, people get fat instead of losing weight, or get skinny instead of gaining weight.
Recently, we launched CalorieCodex (a calorie tracking app with an optional Full Agent Harness, where an agent tracks calories for you) and Symbiote (a programmable workout app; think BoostCamp, but user-programmable so it can run any bodybuilding program).
Yes. The whole time. But, you still need to write good specs if you want well-designed software. Agents are not magic.
- Write good specs into files, usually Markdown.
- Go through the spec files with the agents, have it point out holes and problems, and then update the specs. Go through a few iterations of this and then start your agents on development.
When I'm building something new there's a good chance I don't know about a lot of problems (unknown unknowns). How do you overcome this with just spec files?
I keep coming back to this and the only way to develop properly without tons of hacks patching bugs afterwards is by working at a low level. I don't really write code but I work at the code level still and have Claude write each function and whatnot.
For personal projects, I've rewritten bounded integration parts for things all the time, and so far it's been fine. Coding agents are RL'd to get the thing done, they'd rather get to 95%, not realize there's some fundamental architectural flaw, and hack the last 5%, rather than taking that lesson and redesigning. That's your job.
For business settings, yeah I dunno, a ton of rewriting all the time is not seen as a good look.
The way I have been doing it is to use LLMs to generate the code that I don't want to write: prototypes, tests, benchmarks, isolated,straightforward almost copy-paste code. I still write my own code as before because I enjoy doing that and because trying to understand and fix what an LLM generates and regenerates is harder and more tedious and time consuming than writing the code the way I want to do it in the first place.
What's the problem of the solution proposed ? Don't read/write code anymore. Have strong harness. That's how my team of ~30 has been operating for the most part.
The problem we are trying to solve was never to write code, was to solve business problems
Why are you so aggressive ? What I am saying is how many folks are working nowadays and DHH gives an example of this in his linux distro and in his interview talking about his products.
Maybe try it out with an open mind ?
Do you expect me to break my work contract to convince a stranger on the internet that they're wrong ? lol
I know everything about Omarchy and latest vibe coded version just sucks like many’s have underlined already. I’m open minded to the point that I’d pay to look at your work. It’s always who can’t break the contract that we should trust, no one working in the open, and in the meanwhile software sucks more and more
You are not open minded, you sound like you dont even have a mind of your own. You want to name and shame, just like the rest of your tribe, and because you can't you are just throwing toys out of the pram. Traditionally, I would tell you to grow up, but in your case that would mean to be reborn a few hundred times.
Why would I work on the open tho ? The reason you always get that answer is because it's the norm. I am paid to develop software for a business like the majority of people I suppose.
Outside of my work I spend time with my family as I stare at screens enough during the day.
Maybe if you asked 20 years ago I'd be able to share some of my passion projects, unforutnately AI didn't exist back then
"I don't understand the code anymore. It works." That's fine... today.
"I will never need to understand the code again" is a much different statement.
If you don't understand the code, and the code wasn't written by any human, when you're eventually painted into a corner, how hard is it going to be to get out? Will it be easier or harder than maintaining an understanding through however long it takes to get to that point?
You may be betting your company on the answer. How sure are you?
How is this different than working at a company with hundreds of developers, on a massive codebase? Do you ever really understand the entirety of that?
No. But somebody understood the high-level architecture, and somebody understood each piece in detail. And if you needed to work on piece X.Y.Z, you could find somebody who could tell you how X fit into the big picture, and somebody else who could tell you how Y fit in X, and somebody else who could tell you how Z fit into Y, and maybe somebody else who could tell you about the details of Z.
Its a spectrum. For ai-maintained code (like a gui to visualize performance data), IDGAF what the code looks like. I just let claude or codex go nuts and 100% vibe code.
For code I care about, I audit every single hunk as its produced. I give it extensive style guidelines, and crack down on things like a 20-line essay in a comment.
For mission critical code, I write the code myself and have an agent review it.
After vibing myself into a corner multiple times on important projects, I now have only two modes: clankermaxx for code I don't really care about (mostly frontend react), and write by hand everything else. Using Django for backend already removes the most of the cruft, and writing by hand also means I actually understand what's going on. Works fine for now.
I'm still on the fence for tests: I don't really like to write them, but the LLM-generated tests are pretty bad, even the frontier models on xhigh thinking. I usually generate them, but I don't really have the confidence they test anything except 1==1. Unfortunately it's hard to justify the time spent on writing them manually.
I'm really happy with my opencode + open weight setup, the code is generally pretty good, but I do spend tokens having agents go look for common ai slop patterns.
It's heavily customized, replaced most internal systems via plugins, a set of custom agent instead of builtin ones, different model families for different sub tasks. (don't have claude review its own code)
I'm working on polishing them up and porting a few more from my own harness, then will be open sourcing. Keep your eye out for a "better-opencode" plugin suite, I'll be sure to share it with HN :]
In the near-term, I really like GLM 5.3 prose for code explore / review, give it a shot, it's cheaper (flash model) and catches all sorts of mistakes from the Big Ai models. We fully rolled out our custom pr-review on glm-5.3-flash last week. Most devs are still on claude, moving them towards fireworks and opencode.
OpenCode Go is a great way to try out open weight models for $10/month
I think an underappreciated aspect is that we previously amortized the reading and comprehending over the time we spent designing and writing. I still find myself looking at changes expecting that I can review it in a few minutes when I am in effect coming in cold. That part of the job has changed, but fwiw it isn't so different than the experience I had reviewing another team's changes when I didn't know their services.
I have found the right harness helps, if you have a setup that gives you memory, specs and a graph-based understanding (that part is key) you can be very productive. For me, that's because being able to understand the big picture helps, the challenge of reading the code is always there but starting with 'this method/class/etc does X, so that means this is doing Y..." makes it easier and lets you come in warmer.
I generally produce nearly the same code I'd write myself about 5x faster with AI. I don't just let Claude Code run wild for a long time and have a mess to review. I have it do small chunks I can quickly review, give it feedback, iterate, etc until I like the output, then I move on to the next step. This takes time of course but I've found it faster than reviewing a giant mess of a code review.
My experience as well. Opus 5.5 is really good at matching the existing coding style. I often look at code and think I wrote it, only to realize it was written by Opus.
The trick is to give it enough guidance about the style you prefer.
Share software you’re building
Share the concrete problems you have with LLMs and we might be able to help you out.
if you were serious you would know this is an unserious bar to try and set.
Why? It would help understand what “the same code I would’ve written” means.
We have multiple apps, but I'll share only the free ones; the others are B2B (they haul in major revenue for us, and businesses dealing with us will not like us disclosing this here), largest free one is MacroCodex with 17,000+ users.
We do not have this problem; we heavily use the models from OpenAI and Anthropic, but we have a custom harness built for ultimate cache efficiency and a purpose tuned workflow.
Codex is GREAT but it keeps updating and its UI changes often (remember mini git sidebar ui?)
Our Custom Harness does multiple things, like having a "Skill Selector," which uses our Gambler v1 (26B decision model responds in <100ms locally) and overlays tools on top of existing ones (as the models are RLVL over the specific tool in their harness, we overlay those tools with a familiar interface or format and inject custom features we want the model to interpret). Other than this, it leverages local models to accelerate development. Also, instead of providing custom tools, we actually use the same tools that their original model-provided harness uses. We've a list of 100s of such optimizations; these are just a few from the top of my head.
I have been working on harnesses for a very long time; I read research papers and implement them.
The apps use Flutter and Rust for all of our apps; no ads, no subscriptions.
So, macrocodex is basically a neural network product which "predicts" maintenance calories at a very high accuracy, https://macrocodex.app/
MacroCodex beats everything in its category and it does it for free, no ads, no subscriptions.
Symbiote and CalorieCodex will soon follow the suite.
If it fails, people get fat instead of losing weight, or get skinny instead of gaining weight.
Recently, we launched CalorieCodex (a calorie tracking app with an optional Full Agent Harness, where an agent tracks calories for you) and Symbiote (a programmable workout app; think BoostCamp, but user-programmable so it can run any bodybuilding program).
You can find some screenshots here: https://macrocodex.app/guides/peak-week/
Yes. The whole time. But, you still need to write good specs if you want well-designed software. Agents are not magic.
- Write good specs into files, usually Markdown.
- Go through the spec files with the agents, have it point out holes and problems, and then update the specs. Go through a few iterations of this and then start your agents on development.
Garbage in, garbage out.
> Write good specs into files, usually Markdown.
When I'm building something new there's a good chance I don't know about a lot of problems (unknown unknowns). How do you overcome this with just spec files?
I keep coming back to this and the only way to develop properly without tons of hacks patching bugs afterwards is by working at a low level. I don't really write code but I work at the code level still and have Claude write each function and whatnot.
For personal projects, I've rewritten bounded integration parts for things all the time, and so far it's been fine. Coding agents are RL'd to get the thing done, they'd rather get to 95%, not realize there's some fundamental architectural flaw, and hack the last 5%, rather than taking that lesson and redesigning. That's your job.
For business settings, yeah I dunno, a ton of rewriting all the time is not seen as a good look.
> Write good specs into files, usually Markdown.
Specs that gets systematically ignored way too often XD
Are you still saving time or effort using coding agents this way? How much (estimated)?
The way I have been doing it is to use LLMs to generate the code that I don't want to write: prototypes, tests, benchmarks, isolated,straightforward almost copy-paste code. I still write my own code as before because I enjoy doing that and because trying to understand and fix what an LLM generates and regenerates is harder and more tedious and time consuming than writing the code the way I want to do it in the first place.
What's the problem of the solution proposed ? Don't read/write code anymore. Have strong harness. That's how my team of ~30 has been operating for the most part.
The problem we are trying to solve was never to write code, was to solve business problems
Share the software you’re developing or you’re just a shiller
What is a shiller ? lol
You can look at people who are building on the public like this, for example, DHH with https://omarchy.org/ https://www.youtube.com/watch?v=NYFGCESmikA
I work in private software that I cannot share, but it's for a company you have heard of
I don’t care about DHH I never needed any of his products. I want to see yours
> I work in private software that I cannot share, but it's for a company you have heard of
Sure it’s always like that
Why are you so aggressive ? What I am saying is how many folks are working nowadays and DHH gives an example of this in his linux distro and in his interview talking about his products.
Maybe try it out with an open mind ?
Do you expect me to break my work contract to convince a stranger on the internet that they're wrong ? lol
> Maybe try it out with an open mind ?
I know everything about Omarchy and latest vibe coded version just sucks like many’s have underlined already. I’m open minded to the point that I’d pay to look at your work. It’s always who can’t break the contract that we should trust, no one working in the open, and in the meanwhile software sucks more and more
You are not open minded, you sound like you dont even have a mind of your own. You want to name and shame, just like the rest of your tribe, and because you can't you are just throwing toys out of the pram. Traditionally, I would tell you to grow up, but in your case that would mean to be reborn a few hundred times.
Why would I work on the open tho ? The reason you always get that answer is because it's the norm. I am paid to develop software for a business like the majority of people I suppose.
Outside of my work I spend time with my family as I stare at screens enough during the day.
Maybe if you asked 20 years ago I'd be able to share some of my passion projects, unforutnately AI didn't exist back then
Okay. Don’t share the code, just show us the product.
"I don't understand the code anymore. It works." That's fine... today.
"I will never need to understand the code again" is a much different statement.
If you don't understand the code, and the code wasn't written by any human, when you're eventually painted into a corner, how hard is it going to be to get out? Will it be easier or harder than maintaining an understanding through however long it takes to get to that point?
You may be betting your company on the answer. How sure are you?
How is this different than working at a company with hundreds of developers, on a massive codebase? Do you ever really understand the entirety of that?
No. But somebody understood the high-level architecture, and somebody understood each piece in detail. And if you needed to work on piece X.Y.Z, you could find somebody who could tell you how X fit into the big picture, and somebody else who could tell you how Y fit in X, and somebody else who could tell you how Z fit into Y, and maybe somebody else who could tell you about the details of Z.
Not really, no.
People come and go, documentation isn't great.
It's no different.
If I need to understand complex codebase with thousands of humans touching nowadays I am asking AI anyways
Its a spectrum. For ai-maintained code (like a gui to visualize performance data), IDGAF what the code looks like. I just let claude or codex go nuts and 100% vibe code.
For code I care about, I audit every single hunk as its produced. I give it extensive style guidelines, and crack down on things like a 20-line essay in a comment.
For mission critical code, I write the code myself and have an agent review it.
After vibing myself into a corner multiple times on important projects, I now have only two modes: clankermaxx for code I don't really care about (mostly frontend react), and write by hand everything else. Using Django for backend already removes the most of the cruft, and writing by hand also means I actually understand what's going on. Works fine for now.
I'm still on the fence for tests: I don't really like to write them, but the LLM-generated tests are pretty bad, even the frontier models on xhigh thinking. I usually generate them, but I don't really have the confidence they test anything except 1==1. Unfortunately it's hard to justify the time spent on writing them manually.
Fable 5.1 is better than most engineers if prompted correctly
Step one: just because claude can generate +50k/-50k line PRs doesn’t make it good.
Optimize the PR flow first.
I'm really happy with my opencode + open weight setup, the code is generally pretty good, but I do spend tokens having agents go look for common ai slop patterns.
It's heavily customized, replaced most internal systems via plugins, a set of custom agent instead of builtin ones, different model families for different sub tasks. (don't have claude review its own code)
I'm working on polishing them up and porting a few more from my own harness, then will be open sourcing. Keep your eye out for a "better-opencode" plugin suite, I'll be sure to share it with HN :]
In the near-term, I really like GLM 5.3 prose for code explore / review, give it a shot, it's cheaper (flash model) and catches all sorts of mistakes from the Big Ai models. We fully rolled out our custom pr-review on glm-5.3-flash last week. Most devs are still on claude, moving them towards fireworks and opencode.
OpenCode Go is a great way to try out open weight models for $10/month