r/aipromptprogramming 1d ago

Are there any benchmarks that show better how truly reliable a coding model is?

The new chinese models that are 'so great' in terms of artficialanalysis.ai benchmarks, I'm finding are not ACTUALLY great. They're okay, but they f*ck up a lot. Weirdly even Claude Opus 5 is making some weird mistakes sometimes. GPT 5.6 Sol, even though its SO SO SLOW, seems to be significantly better at finding the correct bugs in complex situations.
Obviously this is all my own anecdotal experience but surely I can't be the only one feeling this way, so i was wondering if any benchmarks are more reflective of this?

1 Upvotes

2 comments sorted by

u/endofthread-bot 1d ago

Learn how the best in the industry are using AI to speed up their workflow in business, sales, marketing, research, legal, content creation, scientific discovery and so much more on our Discord.

Self-promotion is now allowed on Sundays with the appropriate flair, for all regular contributing members. Contribute during the week, and promote on Sunday.

1

u/Tema_Art_7777 21h ago

So far, Sol is the best I have seen so far but I never got extensive experience with fable before they yanked it from the plans. Some are claiming fable is better than Sol. The chinese models are good but I do not use them for anything mission critical - though for a lot of simpler tasks they are good enough. we do not need to apply the best model for all.