pull down to refresh

I see the lack of accountability and slop avalanche as a product of our evaluation bottleneck. Truly think that's the problem that needs to be solved: our ability to actually process the flood of unlimited AI generation we're facing.

I see what you mean about content discovery.

I agree with that. I think that "AI will take your job" made a lot of people overconfident. Also "I don't read, I ship" in January was a terrible example. But I can see a lot of devs liking that because now you don't have to review deeply, or think. Which is the only thing you should be doing. And it's the least fun part!

To some degree, we can fight fire with fire. But you cannot let an LLM be a judge. You cannot ask it: "is this ok?" What you can ask is: "find all the problems." And then you have to read it to make sure that those really are problems. Currently, hit rate is... 90% or so, so I don't use it for critical things because a 10% error rate compounds rapidly, also when these are false positives.

reply