Rendered at 00:32:48 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Johnny_Bonk 55 minutes ago [-]
I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
dhorthy 50 minutes ago [-]
yeah this was just a start - the fastest cheapest thing we could try for a brand new model.
I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark
I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model
Johnny_Bonk 38 minutes ago [-]
yeah i agree
dan_gee 32 minutes ago [-]
As someone who doesn't use AI, this is totally incomprehensible to me.
You might as well be talking about the difference between smoking Maui Wowie and Grandaddy Purp.
scrollaway 27 minutes ago [-]
How is this useful or insightful?
You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
dan_gee 25 minutes ago [-]
My point is that the differences between these models are so minor that obsessively benchmarking them comes across as navel-gazing.
joatmon-snoo 3 minutes ago [-]
The evidence that proves a model is actually a step function change is these benchmarks.
If a model isn’t a step function change? Welcome to research.
Johnny_Bonk 7 minutes ago [-]
like all good science, measure everything
killingtime74 34 minutes ago [-]
Did you not benchmark latest GPT 5.6 or GLM 5.1/Kimi K3 because of cost? I can run them if you share how you ran them
dhorthy 29 minutes ago [-]
no i'm spinning those up at some point this week. here's the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant messages but it should get you kicked off.
> fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 https://arxiv.org/html/2603.24755v1
> Get all the context, fetch any mentioned repos, and then propose a plan to me.
> i have an anthropic API key in ....
> Let's do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let's try it. What do you need from me?
> Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.
> Actually let's do one easy, one medium, and one hard problem please. If we have a hard problem I'd like to see that.
> lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable
knighthacker 9 minutes ago [-]
This is where Opus 5 shines
dcl 57 minutes ago [-]
finally the benchmark for me
dhorthy 48 minutes ago [-]
i hope that is because you hate slop and not because you write it
cute_boi 16 minutes ago [-]
Opus 5 is an overconfident stupid model. It tries to generate too much slop, tries to act like everything will fall. I have reversed back to fable and codex sol.
dhorthy 13 minutes ago [-]
yes sol is still my daily driver for most coding tasks
I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5)
but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8
I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark
I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model
You might as well be talking about the difference between smoking Maui Wowie and Grandaddy Purp.
You ever go to forums full of entomology specialists and tell them you don’t understand their fancy terms?
If a model isn’t a step function change? Welcome to research.
> fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 https://arxiv.org/html/2603.24755v1 > Get all the context, fetch any mentioned repos, and then propose a plan to me.
> i have an anthropic API key in .... > Let's do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let's try it. What do you need from me?
> Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.
> Actually let's do one easy, one medium, and one hard problem please. If we have a hard problem I'd like to see that.
> lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable
I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5)
but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8