How do you measure whether it works better day to day without benchmarks?

bulbar · 2025-12-11T18:48:13 1765478893

Manually labeling answers maybe? There exist a lot of infrastructure built around and as it's heavily used for 2 decades and it's relatively cheap.

That's still benchmarking of course, but not utilizing any of the well known / public ones.

verdverm · 2025-12-11T18:51:21 1765479081

Internal evals, Big AI certainly has good, proprietary training and eval data, it's one reason why their models are better

aydyn · 2025-12-11T18:58:53 1765479533

Then publish the results of those internal evals. Public benchmark saturation isn't an excuse to be un-quantitative.

verdverm · 2025-12-11T19:03:01 1765479781

How would published numbers be useful without knowing what the underlying data being used to test and evaluate them are? They are proprietary for a reason

To think that Anthropic is not being intentional and quantitative in their model building, because they care less for the saturated benchmaxxing, is to miss the forest for the trees

aydyn · 2025-12-11T20:17:53 1765484273

Do you know everything that exists in public benchmarks?

They can give a description of what their metrics are without giving away anything proprietary.

verdverm · 2025-12-11T23:00:54 1765494054

I'd recommend watching Nathan Lambert's video he dropped yesterday on Olmo 3 Thinking. You'll learn there's a lot of places where even descriptions of proprietary testing regimes would give away some secret sauce

Nathan is at Ai2 which is all about open sourcing the process, experience, and learnings along the way

aydyn · 2025-12-12T08:18:17 1765527497

Thanks for the reference I'll check it out. But it doesnt really take away from the point I am making. If a level of description would give away proprietary information, then go one level up to a more vague description. How to describe things to a proper level is more of a social problem than a technical one.

verdverm · 2025-12-13T20:26:13 1765657573

You seem stuck on the idea that they should have to share information when they don't have to. That they share any is a welcome change. Push too hard and they may stop sharing as much

standardUser · 2025-12-11T18:46:04 1765478764

Subscriptions.

mrguyorama · 2025-12-11T19:39:29 1765481969

Ah yes, humans are famously empirical in their behavior and we definitely do not have direct evidence of the "best" sports players being much more likely than the average to be superstitious or do things like wear "lucky underwear" or buy right into scam bracelets that "give you more balance" using a holographic sticker.

standardUser · 2025-12-15T16:42:00 1765816920

It's all the shareholders care about. These are not research institutions.