pull down to refresh

5.2 "felt" better than the benches, interested to give this one a play

Grok 4.6 was feeling a lot better than 4.5 for a few days, fast and not lazy, but yesterday started getting slow and lazy again... I swear Cursor is A/B testing on us constantly.

It's always hit & miss when it's post-training based. I.e. gpt 5.1/5.5 and opus 4.8 all had regressions on functionality out of the training scope (or well, for 4.8 I guess the training wasn't great to begin with.) I'll check it out when it's open.

Funnily, for 5.2 my best harness is plain codex-oss with gpt prompts and they don't even mention that here. So I'll probably have to do a bunch of tests again.

reply