> We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200.
If this holds up that's a really big deal.
wayfwdmachine 1 days ago [-]
Huge if true. As it were.
pvillano 1 days ago [-]
A lot of people are betting their money on infinite growth forever of AI performance, compute usage, user base, subscription price.
I think cost will decrease forever.
bilater 1 days ago [-]
Cost will keep dropping, but the frontier will keep getting pushed. The whole "give me today's model 10x cheaper and I'm good" line is a fallacy. It isn't true now and it never will be for the top 1% of tasks, which will create the most economic gains.
OtherShrezzing 1 days ago [-]
> It isn't true now and it never will be for the top 1% of tasks, which will create the most economic gains
Vulcan Materials Company has produced some of the strongest and most resilient economic gains for investors at around 25% gross margins for 50 years.
Vulcan’s business is crushing rocks, and then driving those rocks to where people need crushed rocks.
I mention it because Vulcan isn’t a sophisticated business at its core, but its economic returns are exceptional, because they do their unsophisticated work exceptionally well, and exceptionally efficiently.
Across the economy, most economic gains are created by companies like Vulcan, who do boring repetitive work exceptionally well and efficiently.
I expect this to hold true into the era of AI. Most stuff probably doesn’t need an exceptional model, and paying for an exceptional model to do unsophisticated work will leave your business vulnerable to competitors who take time to find the most efficient model for the task, and undercut you.
yowlingcat 1 days ago [-]
I just wonder whether the analogy holds from atom based businesses to byte based businesses. At the end of the day, there are pretty hard constraints on how much a competitor could undercut Vulcan - they'd need to be better a) crushing rocks and b) driving those rocks to where people need it. Something about that, while simple, seems really nonfungible to do well. The business may traffic in commodities, but the service and quality it provides ironically seems to be not really a commodity.
Conversely, AI models feel pretty different. People will swap models at the drop of a hat and while not perhaps fully fungible, the cost of switching can be as low as one engineers it to be. In this case, the service does really seem to feel more like a commodity.
Not sure where this leads or how this ends. Perhaps those two will swap poles for me (IE as robotic automation improves, there are more commoditized Vulcans, and as metered intelligence improves, they become less fungible). I just can't quite see how yet. Just food for thought.
adam_arthur 1 days ago [-]
There are an enormous number of tasks that can get by on good enough.
If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out.
And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier.
Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years.
5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain?
The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months.
I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.
bilater 1 days ago [-]
That can all be true but the frontier models will still have a huge market. You're thinking of all tasks as a fixed pie. The top 1% of intelligence opens up a whole new pie, stuff nobody does today because it's too expensive: daily cancer scans instead of one every few years, asteroid mining missions that need ten thousand PhD-hours of planning, custom drugs designed for your specific tumor, a personal lawyer and doctor for every person on earth, auditing every line of code in every bank and hospital continuously and so on.
throwup238 1 days ago [-]
> daily cancer scans instead of one every few years
Silly nitpick: the reason we don’t do daily cancer scans isn’t the cost, it’s the false positive rate. Invasive procedures like biopsies come with complications like infections that happen at a higher rate and do more damage than the cancer that doesn’t even exist. This dilemma is pervasive in medicine, because our tests aren’t perfect but the thing they’re testing for is rare.
sacred_numbers 1 days ago [-]
The whole paradigm changes, though, when you can do daily cancer scans. You don't get a biopsy when the scan shows a lump. You get a biopsy after a couple weeks of daily scans showing the lump growing. Plus, having all the data from the daily scans improves your testing accuracy so false positives and negatives are more rare.
throwup238 1 days ago [-]
The errors are correlated, not random. Lumps are usually benign cysts. If you start cutting people open for every cyst on a scan, you’d kill a lot more people than you’d save.
This isn’t something you can solve with more scans because the tests test for data that is indistinguishable. They look the same on a scan, there’s an overlap in the assay with some random protein with the same binding sites that is only present in 1% of the population, the coding gene in one person gets repeated in a noncoding region in another, and so on. The “more data” that works is a doctor applying professional judgement (which they’re also famously bad at because biology is a fickle mistress).
ACCount37 1 days ago [-]
If you're doing non-redundant tests and your uncertainty bars aren't shrinking, it's usually a skill issue.
If it looks like a duck, it might be a duck - or a painting of one. If it looks like a duck, swims like a duck, and quacks like a duck? The joint duck estimation is much more confident now. There might be a few more observational tests one should administer before committing to a duckhood decision, but each tests pins down variables and rejects confounders. Uncertainties are cut down, and we get closer to crossing the threshold between "duck-informative" and "duck-actionable".
Thus, it's often worth it to improve observability. If you managed to make a certain test more reliable, or cheaper to administer, or reduced the chance of adverse effects? Or, in other words, improved SNR, reduced costs, and reduced costs? You can get more information for your buck. Paired with good knowledge: you can make better decisions more easily.
The fact that the thought of "having more information might be bad actually" even occurs in the field of medicine shows just how far it is from being optimal. Having more information isn't always beneficial - some information is genuinely redundant. Some information is not worth the effort of gathering and integrating it. But if you get more information and it results in worse outcomes? You're doing something wrong.
22 hours ago [-]
15 hours ago [-]
mhluongo 1 days ago [-]
More data (daily scans) can mean we get better at medicine, though, and more accurate. You're assuming "daily cancer scans" look just like they do today, rather than eg unobtrusive devices in our daily lives that measure changes over time.
The frequent "muh false positives" comment we hear from doctors appears to be a lack of imagination?
1 days ago [-]
DeluluDon 23 hours ago [-]
It comes from experience.
adam_arthur 1 days ago [-]
Yes, agree that token consumption will increase exponentially for the next while.
Disagree that the frontier model is where the economic gains will be realized.
The smaller the relative gap between frontier and non-frontier/open weights, the less pricing power.
This gap has shown only to shrink over time, not expand.
Businesses will pay more for frontier, but not meaningfully more to justify the economics. It's always going to be a low margin business, perhaps outside of cyber security, warfare/intelligence and perhaps drug discovery.
Though the expensive and time consuming part of drugs is doing the trials and getting approval, not coming up with ideas
pixl97 1 days ago [-]
Sounds kind of like another K shaped economy. Low end models will be highly competitive and low profit. Problems that can be solved by low end models will be highly competitive and low profit too.
Where the interesting work will be is at the median point where cheap models do almost all of it but need to hand off some parts to the SOTA/more expensive models. Seems like there's money to be made by maximizing low end use while maintaining quality.
adam_arthur 1 days ago [-]
Certainly there's still a business there, I'm not saying they won't exist. But it's not going to be a monopoly-esque business with so many players in the ring, OpenAI, Anthropic, Google, Meta, Deepseek, Alibaba, GLM, Kimi etc. It will be cutthroat and a race to the bottom on price. And the difference from today -> 6 months ago intelligence will not be very meaningful.
Investors are largely treating these as future monopolies though.
We can already do so much with existing models. Harness improvements are probably more meaningful at this point.
e.g. say most image recognition can get saturated by a model of size xB parameters, so your tool for that can handoff to a smaller model. Document text extraction can use a model of size yB parameters. A model of size zB for summarizing text.
We are starting to get to a point where you can reasonably scope out an upper bound of required size/effort for many common tasks, and if you string these together, the frontier will largely act as an intelligent invoker of more efficient models.
Up until now there have been meaningful gains to each of those types of workstreams by using newer models, but that is starting to no longer be the case.
Yes, I do believe token consumption will rise exponentially from here in the near term. But cost of switching is low, and substantial profitability will be difficult.
bilater 1 days ago [-]
how much would you pay for a prompt that could cure cancer? if you're a pharma company you would pay millions to get there days faster than your competitor. as intelligence rises the marginal value it can deliver rises with it.
philipkglass 1 days ago [-]
Something like curing cancer (more realistically, curing a specific kind of cancer) has to interact with much slower real-world processes. The most expensive part of drug development is Phase 3 clinical trials in humans. Even the smartest model in the world can't accelerate that meaningfully. Even much earlier when drugs are just testing in cell cultures, it's a lot slower to run lab tests than to run software tests or mathematical proof checkers.
Or to put it another way, there's enough natural variation in real-world bottlenecks that no pharma company can assume they'll beat competitors to market by using a smarter model.
A really smart model could significantly improve the pharma business if it could identify promising approaches to cancer treatment that are less likely to fail in clinical trials, but I don't think that the frontier labs have data to make that work yet. Much of the biomedical literature is poorly reproducible ("replication crisis") and much of the drug-development-specific data is proprietary, never published in the first place.
I do have hopes that general laboratory automation will go faster with LLM assistance, even if all the LLM does is write Python glue scripts to enable custom workflows and instrument integrations.
done_lurking 23 hours ago [-]
Wouldn't daily cancer scans give you cancer from all the radiation? I thought that's why the doc hides behind a wall when giving an X-ray.
michaellee8 1 days ago [-]
yea you see no body are vibecoding games before opus 5 and astra, after they are released games basically got commoditized
adam_arthur 1 days ago [-]
If the output is commoditized, how much can you afford to pay for the input?
sampullman 1 days ago [-]
Is this sarcasm? I assume most studios are integrating AI into their workflows, but I still haven't seen a single vibecoded game that looks interesting.
chasd00 24 hours ago [-]
I was yelled at repeatedly here that AI is not used in gaming.
saulpw 1 days ago [-]
I think there's an intelligence limit, or at least asymptote. It may be above human intelligence, but I don't think it's miles above it (at least not the kind of intelligence humans can create, recognize, or use). For example in Go, most estimates place God or "perfect play" three ranks above top professionals[0]. In the latest human-AI Go match, the human got a 2 stone handicap. So it's not like we have a lot more frontier to push there.
Just FTR, Shin Jinseo's result in that latest match is seen as extraordinary. In the last few years it's been common for top pros to play bots on 3 or even 4 stones.
The thing about playing strength in games like go is that it doesn't map at all linearly to results. It could very well be the case that the bots are only a few points of handicap away from perfect play, but thousands of points of ELO.
Anyway, this is a poor argument for a possible "limit to intelligence" because of course the game is finite (even if the search space is extraordinarily large). Of course you can't beat the "hand of god", by definition. Of course you can create games and puzzles for which "hand of god" is a coherent concept. It shouldn't be too hard to specify games where it isn't, though (perhaps something involving real numbers).
pixl97 1 days ago [-]
Intelligence is spiky. In some things humans may play near the limit (Go possibly), but if you look at parts of mathematics like addition, humans can add in their head just fine but its a few trillion times faster to use a computer on addition problems of any size.
And that's not even really touching societal/network intelligence. A single human isn't that smart and can't accomplish that much. Hence we form families, and companies, and societies, and governments. What does a society of AIs look like?
saulpw 1 days ago [-]
Note that addition that's "trillion times faster" is purely an optimization, not greater intelligence.
And a society of humans does not increase our overall intelligence. It allows all of humanity to access the accomplishments of our most intelligent members throughout history (which is huge, don't get me wrong!) but it seems almost self-evident that our civilization as a whole is not smarter than Newton/Einstein/von Neumann/etc.
pixl97 1 days ago [-]
Intelligence is also access to data, along with all of it's other requirements. Society as a whole is far smarter than Einstein.
If I dump Einstein on an island at 5 you don't get a theory of relativity. The data provided by a civilization is required, hence why people 10,000 years ago didn't go to the stars. It is unlikely they were significantly less intelligent than us. They just didn't have the language of data we do now.
bilater 1 days ago [-]
this assumes our whole universe and what we can do in it is a finite go board. maybe it is. but we are no where close to exploring even a fraction of it. lots of things need to happen before any limit is reached. the cavemen would also probably think we saturated tools once they saw bows and arrows.
czhu12 1 days ago [-]
its already true today according to openrouter usage stats. https://openrouter.ai/rankings Most people just use cheaper open weight
kennywinker 1 days ago [-]
Intuitively I agree with you, but you've cited evidence that doesn't actually prove what you said. Yes, openrouter users choose cheaper open weights. If you're on openrouter, you're already shopping around. We'd need stats for claude, codex, gemini, grok (yuck), AND openrouter, to determine what people in general (not just the subset of people who decided to shop around) are actually doing.
czhu12 1 days ago [-]
"Almost nobody is using fable 5" at anthropic apparently
I agree, but I also think as AI gets better we're going to see Jevons paradox in full swing, which might delay lower costs. We've seen this with Astra according to Tibo: https://x.com/thsottiaux/status/2097559315150426222
ansk 1 days ago [-]
I don't know enough about the specific models they're comparing against to say this definitively, but it looks to me like they're comparing their pre-trained models with others' post-trained models.
The metric upon which their 10x claim is based (bits-per-byte) is exactly the metric which is optimized during pre-training. Post-trained models are fine-tuned to optimize other metrics, which is known to be detrimental to performance on bits-per-byte evaluations. So bits-per-byte evaluations will always make a pre-trained model look favorable in comparison to a comparable model which has also undergone post-training.
Can someone confirm whether the models they are comparing against (DeepSeek V4, Kimi K2, and Nemotron 3 Ultra) have been post-trained?
KaushikR2 22 hours ago [-]
See footnote 2:
> “Base models” are pretrained models that have not yet undergone reinforcement learning, SFT, or other post-training. They are highly sensitive to prompting, making sampling-based evals unreliable. Instead, we measured bits-per-byte loss on heldout data, which does not suffer from prompt sensitivity and smooths measurement of otherwise emergent abilities. As a side note, we were surprised that Nemotron 3 outperforms DeepSeek V4 Pro across the board but found this to be consistent across domains and inference engines. This might indicate that Nemotron’s weak performance on benchmarks after RL is due to weaker post-training, but the pretrain was ahead of Chinese open-weight competitors.
brrrrrm 1 days ago [-]
they say they're looking at base models, so I think it's fairly compared as written.
ismael_rr 1 days ago [-]
Super awesome. Wish they would release the paper about what they did to achieve this. I remember nous released the token superposition paper which improved pretraining FLOPs some, but not 50x: https://nousresearch.com/token-superposition. Wondering if they also found some cool tokenization strategiesa
pixl97 1 days ago [-]
"Write a paper" < "Sell to a big AI lab for $$$"
Going to be interesting to see what happens to discoveries like this in the future.
brrrrrm 1 days ago [-]
this is basically the only thing pre-training teams work on in labs. compute efficiency is the metric, the assumption that scaling = intelligence is considered a given.
monneyboi 1 days ago [-]
Imagine the sheer amount of power you could save by releasing the paper.
speedgoose 1 days ago [-]
But thanks to the Jevon Paradox, the global power consumption would probably increase.
Solar is now not only the cheapest energy it's cheaper to intitially deploy than non-renewables.
Doesn't mean we should waste energy; it does mean that we have crossed a threshold beyond which energy concerns change shape.
pixl97 1 days ago [-]
"Honey, why is there a solar panel-maximizer converting our car?"
pvillano 1 days ago [-]
Sunlight no longer reaches the earth's surface but at least we know P vs NP
corysama 1 days ago [-]
You just need to work your way down to a deeper layer in the https://en.wikipedia.org/wiki/Matrioshka_brain Not the deepest layer. Humans would fry instantly there. But, the warm glow of the third layer down upon the fourth is quite pleasant.
pixl97 1 days ago [-]
I mean P=NP might be a fair trade.
vatsachak 1 days ago [-]
Cool story. If it's true the company will be bought by open AI/Anthropic and Chinese labs will discover the trick and open source it by next quarter.
vkaku 1 days ago [-]
Next Quarter? :) That's too long
gdiamos 1 days ago [-]
Training improvements are very easy to copy.
vkaku 1 days ago [-]
This is great. All algorithmic efficiencies are amazing!
One thing I'd remind all scientists and the wonderful people here is this wonderful meme/line from Jurassic Park: "Your scientists were so preoccupied with whether they could they didn't stop to think if they should."
What is the actual amount of data that needs to be pre-trained and what is not? Nobody has come up with great answers to this question, and I'm already seeing amazing 0.5b-2b parameter models working very well with n-Gram corpuses of data. So, how many parameters do you really need for a given workload?
riazrizvi 1 days ago [-]
Very helpful thanks.
mohsen1 1 days ago [-]
[dead]
FailMore 1 days ago [-]
[flagged]
asadm 1 days ago [-]
not a good enough submarine
pvillano 1 days ago [-]
Imagine yourself the CEO of a big AI company. It takes about a month to develop and train a model, so you release a new model every month. A startup says they can 10x your efficiency. What does that get you? You can't release a new model every three days. You can't 10x R&D either. You definitely can't tell investors that you are growing at the same rate, but selling off assets and cancelling purchasing contracts. So you just never improve efficiency enough to use less energy than the previous model version.
I don't believe this is actually happening.
alex_duf 1 days ago [-]
A 10x reduction in pre-training means a 10x faster feedback loop. I'm pretty sure any lab would sign up for that. You can start experimenting on different approaches much more aggressively.
awestroke 1 days ago [-]
> What does that get you?
Cheaper model training runs? Ability to scale training to larger model sizes without extending training time?
pvillano 1 days ago [-]
Yes, but a 10x larger model is only marginally better, and 10x cheaper training runs is only useful if you can find a use for 10x as many.
aaronblohowiak 1 days ago [-]
more experiments.
impossiblefork 1 days ago [-]
So you can make a huge internal model that you can then distill from?
BoorishBears 1 days ago [-]
This seems weirdly pessemistic: frontier labs have much stronger pretraining than most open weights models
And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference
They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained.
That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way.
This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)
enzyme1234 1 days ago [-]
the obvious answer is that you can try more experiments over the same period of time, so you find more improvements per month, and the rate of improvement of the models you release increases
[1]: https://magic.dev/blog/100m-token-context-windows (also linked to in their blogpost)
If this holds up that's a really big deal.
I think cost will decrease forever.
Vulcan Materials Company has produced some of the strongest and most resilient economic gains for investors at around 25% gross margins for 50 years.
Vulcan’s business is crushing rocks, and then driving those rocks to where people need crushed rocks.
I mention it because Vulcan isn’t a sophisticated business at its core, but its economic returns are exceptional, because they do their unsophisticated work exceptionally well, and exceptionally efficiently.
Across the economy, most economic gains are created by companies like Vulcan, who do boring repetitive work exceptionally well and efficiently.
I expect this to hold true into the era of AI. Most stuff probably doesn’t need an exceptional model, and paying for an exceptional model to do unsophisticated work will leave your business vulnerable to competitors who take time to find the most efficient model for the task, and undercut you.
Conversely, AI models feel pretty different. People will swap models at the drop of a hat and while not perhaps fully fungible, the cost of switching can be as low as one engineers it to be. In this case, the service does really seem to feel more like a commodity.
Not sure where this leads or how this ends. Perhaps those two will swap poles for me (IE as robotic automation improves, there are more commoditized Vulcans, and as metered intelligence improves, they become less fungible). I just can't quite see how yet. Just food for thought.
If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out.
And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier.
Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years.
5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain?
The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months.
I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.
Silly nitpick: the reason we don’t do daily cancer scans isn’t the cost, it’s the false positive rate. Invasive procedures like biopsies come with complications like infections that happen at a higher rate and do more damage than the cancer that doesn’t even exist. This dilemma is pervasive in medicine, because our tests aren’t perfect but the thing they’re testing for is rare.
This isn’t something you can solve with more scans because the tests test for data that is indistinguishable. They look the same on a scan, there’s an overlap in the assay with some random protein with the same binding sites that is only present in 1% of the population, the coding gene in one person gets repeated in a noncoding region in another, and so on. The “more data” that works is a doctor applying professional judgement (which they’re also famously bad at because biology is a fickle mistress).
If it looks like a duck, it might be a duck - or a painting of one. If it looks like a duck, swims like a duck, and quacks like a duck? The joint duck estimation is much more confident now. There might be a few more observational tests one should administer before committing to a duckhood decision, but each tests pins down variables and rejects confounders. Uncertainties are cut down, and we get closer to crossing the threshold between "duck-informative" and "duck-actionable".
Thus, it's often worth it to improve observability. If you managed to make a certain test more reliable, or cheaper to administer, or reduced the chance of adverse effects? Or, in other words, improved SNR, reduced costs, and reduced costs? You can get more information for your buck. Paired with good knowledge: you can make better decisions more easily.
The fact that the thought of "having more information might be bad actually" even occurs in the field of medicine shows just how far it is from being optimal. Having more information isn't always beneficial - some information is genuinely redundant. Some information is not worth the effort of gathering and integrating it. But if you get more information and it results in worse outcomes? You're doing something wrong.
The frequent "muh false positives" comment we hear from doctors appears to be a lack of imagination?
Disagree that the frontier model is where the economic gains will be realized.
The smaller the relative gap between frontier and non-frontier/open weights, the less pricing power.
This gap has shown only to shrink over time, not expand.
Businesses will pay more for frontier, but not meaningfully more to justify the economics. It's always going to be a low margin business, perhaps outside of cyber security, warfare/intelligence and perhaps drug discovery.
Though the expensive and time consuming part of drugs is doing the trials and getting approval, not coming up with ideas
Where the interesting work will be is at the median point where cheap models do almost all of it but need to hand off some parts to the SOTA/more expensive models. Seems like there's money to be made by maximizing low end use while maintaining quality.
Investors are largely treating these as future monopolies though.
We can already do so much with existing models. Harness improvements are probably more meaningful at this point.
e.g. say most image recognition can get saturated by a model of size xB parameters, so your tool for that can handoff to a smaller model. Document text extraction can use a model of size yB parameters. A model of size zB for summarizing text.
We are starting to get to a point where you can reasonably scope out an upper bound of required size/effort for many common tasks, and if you string these together, the frontier will largely act as an intelligent invoker of more efficient models.
Up until now there have been meaningful gains to each of those types of workstreams by using newer models, but that is starting to no longer be the case.
Yes, I do believe token consumption will rise exponentially from here in the near term. But cost of switching is low, and substantial profitability will be difficult.
Or to put it another way, there's enough natural variation in real-world bottlenecks that no pharma company can assume they'll beat competitors to market by using a smarter model.
A really smart model could significantly improve the pharma business if it could identify promising approaches to cancer treatment that are less likely to fail in clinical trials, but I don't think that the frontier labs have data to make that work yet. Much of the biomedical literature is poorly reproducible ("replication crisis") and much of the drug-development-specific data is proprietary, never published in the first place.
I do have hopes that general laboratory automation will go faster with LLM assistance, even if all the LLM does is write Python glue scripts to enable custom workflows and instrument integrations.
[0]https://senseis.xmp.net/?HandOfGod
The thing about playing strength in games like go is that it doesn't map at all linearly to results. It could very well be the case that the bots are only a few points of handicap away from perfect play, but thousands of points of ELO.
I wrote about all of this on Stack Exchange, along with some napkin math, a couple years back: https://boardgames.stackexchange.com/a/61058
Anyway, this is a poor argument for a possible "limit to intelligence" because of course the game is finite (even if the search space is extraordinarily large). Of course you can't beat the "hand of god", by definition. Of course you can create games and puzzles for which "hand of god" is a coherent concept. It shouldn't be too hard to specify games where it isn't, though (perhaps something involving real numbers).
And that's not even really touching societal/network intelligence. A single human isn't that smart and can't accomplish that much. Hence we form families, and companies, and societies, and governments. What does a society of AIs look like?
And a society of humans does not increase our overall intelligence. It allows all of humanity to access the accomplishments of our most intelligent members throughout history (which is huge, don't get me wrong!) but it seems almost self-evident that our civilization as a whole is not smarter than Newton/Einstein/von Neumann/etc.
If I dump Einstein on an island at 5 you don't get a theory of relativity. The data provided by a civilization is required, hence why people 10,000 years ago didn't go to the stars. It is unlikely they were significantly less intelligent than us. They just didn't have the language of data we do now.
https://analyticsindiamag.com/ai-features/almost-nobody-is-u...
The metric upon which their 10x claim is based (bits-per-byte) is exactly the metric which is optimized during pre-training. Post-trained models are fine-tuned to optimize other metrics, which is known to be detrimental to performance on bits-per-byte evaluations. So bits-per-byte evaluations will always make a pre-trained model look favorable in comparison to a comparable model which has also undergone post-training.
Can someone confirm whether the models they are comparing against (DeepSeek V4, Kimi K2, and Nemotron 3 Ultra) have been post-trained?
> “Base models” are pretrained models that have not yet undergone reinforcement learning, SFT, or other post-training. They are highly sensitive to prompting, making sampling-based evals unreliable. Instead, we measured bits-per-byte loss on heldout data, which does not suffer from prompt sensitivity and smooths measurement of otherwise emergent abilities. As a side note, we were surprised that Nemotron 3 outperforms DeepSeek V4 Pro across the board but found this to be consistent across domains and inference engines. This might indicate that Nemotron’s weak performance on benchmarks after RL is due to weaker post-training, but the pretrain was ahead of Chinese open-weight competitors.
Going to be interesting to see what happens to discoveries like this in the future.
https://en.wikipedia.org/wiki/Jevons_paradox
Doesn't mean we should waste energy; it does mean that we have crossed a threshold beyond which energy concerns change shape.
One thing I'd remind all scientists and the wonderful people here is this wonderful meme/line from Jurassic Park: "Your scientists were so preoccupied with whether they could they didn't stop to think if they should."
What is the actual amount of data that needs to be pre-trained and what is not? Nobody has come up with great answers to this question, and I'm already seeing amazing 0.5b-2b parameter models working very well with n-Gram corpuses of data. So, how many parameters do you really need for a given workload?
I don't believe this is actually happening.
Cheaper model training runs? Ability to scale training to larger model sizes without extending training time?
And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference
They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained.
That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way.
This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)