I gave 4 coding agents a $100 budget to build a functional PDF editor. The implementations are okay, but not great: within a few clicks I find bugs in almost all of them. I hypothesize that this is partly because agents don’t interact with software like humans: they don’t pursue concrete goals and experience friction on their way, and they also just look at the UI much less. In consequence, unless we change the way LLMs operate, we will be cursed with annoying SlopWare.
I think the biggest difference is how the feedback loop used to work when hand coding. The pdf colors being inverted in dark mode is a bug i could see myself accidentally create, but i would've seen that immediately when testing if dark mode works at all
I appreciate a more nuanced take than the typical "it's all slop" vs "you're an idiot for handcoding anything anymore" crowds on this site. Especially one backed by an experiment.
My experience is agents really need cheap easily runnable third party ways to validate their work and stay on track. This usually means compilers and test suites but it can also mean visual checks (like all those vibecoded game ports flooding feeds right now). Of course, these things are useful for humans too. Tightening the feedback loop is always good. But for agents, the lack of feedback can be catastrophic. Better models can ward off catastrophe for longer but eventually without enough feedback, the errors not only accumulate but increase the chance of future errors.
To the points in the conclusion, humans still have the edge on dealing with vagueness and lack of clarity in goals. And we fill in the details of our goals in real time. In other words, the act of writing code feeds back into the other parts of building a program. Agents aren't like this, if they can see where they're going they can get there with remarkable tenacity but without it they'll get somewhere goal shaped fast. You can use agents to take big steps for you, but you're going to lose some fine touch.
I gave 4 coding agents a $100 budget to build a functional PDF editor. The implementations are okay, but not great: within a few clicks I find bugs in almost all of them. I hypothesize that this is partly because agents don’t interact with software like humans: they don’t pursue concrete goals and experience friction on their way, and they also just look at the UI much less. In consequence, unless we change the way LLMs operate, we will be cursed with annoying SlopWare.
I think the biggest difference is how the feedback loop used to work when hand coding. The pdf colors being inverted in dark mode is a bug i could see myself accidentally create, but i would've seen that immediately when testing if dark mode works at all
I appreciate a more nuanced take than the typical "it's all slop" vs "you're an idiot for handcoding anything anymore" crowds on this site. Especially one backed by an experiment.
My experience is agents really need cheap easily runnable third party ways to validate their work and stay on track. This usually means compilers and test suites but it can also mean visual checks (like all those vibecoded game ports flooding feeds right now). Of course, these things are useful for humans too. Tightening the feedback loop is always good. But for agents, the lack of feedback can be catastrophic. Better models can ward off catastrophe for longer but eventually without enough feedback, the errors not only accumulate but increase the chance of future errors.
To the points in the conclusion, humans still have the edge on dealing with vagueness and lack of clarity in goals. And we fill in the details of our goals in real time. In other words, the act of writing code feeds back into the other parts of building a program. Agents aren't like this, if they can see where they're going they can get there with remarkable tenacity but without it they'll get somewhere goal shaped fast. You can use agents to take big steps for you, but you're going to lose some fine touch.
Best combo for me: Jev for simple steps, and Fable or Opus only for planning or review
[flagged]
Why reinvent the wheel ?
What new value add are you attempting to create ?
did you read it, no.