The full insight mine on the AI data category and the people building their own models, run through Phase 1 (Input Intelligence) and Phase 2 (Insight Discovery). No strategy, no concepts. Just the truths, the clichés and the nerve, ready to brief from. Focused on DatologyAI, ahead of the launch.
Everyone is racing to build a smarter model. Almost nobody wants to admit the model was never the hard part. What you feed it is.
Models are what they eat. For three years the industry's answer to every problem was more: more data, more parameters, more GPUs. But the high-quality internet is nearly picked clean, compute is rationed, and pouring more of it into junk data just makes a model expensive, not smart. The biggest lever in AI is the least glamorous one, the quality of what goes in, and it is the one thing almost nobody is doing well. Refine that, and ten people can match a lab that raised two billion dollars.
It attacks the category's deepest reflex, which is to worship scale. Bigger models, more tokens, another rack of GPUs. But the buyer is not really trying to have the most. They are trying to build a model that is unmistakably theirs and good enough to bet the company on. The truth quietly demotes the thing everyone obsesses over, the model, and promotes the thing nobody wants to touch, the data. It reframes the whole purchase from an arms race you cannot win into a lever you already hold: the data you own and the quality you put into it.
Every vendor assumes the builder wants a bigger, faster way to train. The builder actually wants to stop being at the mercy of scale, to own intelligence that is theirs, on data nobody else has, without handing it to the one company most able to replace them.
11pm. A founder reads that a two-year-old startup just matched a frontier lab on one benchmark, with a tenth of the team and a fraction of the compute. They think about the GPU invoice they signed last quarter to brute-force past a model that still is not good enough. They think about the proprietary data sitting in their cloud, the thing that is actually theirs, that they have never properly used. And the quiet thought arrives: we have been trying to buy our way out of a problem money does not fix.
The emotional architecture, not concepts. Five separate nerves inside the same person: the founder or head of AI betting the company on a model they build themselves.
A two-year-old with one model publicly outclasses your whole suite on the thing you are known for. On Twitter, in front of your board and your customers, you are suddenly the old guard.
Your model underdelivered and you told the room it was a GPU limit. You never fully audited the data, partly because you were afraid of what you would find. It was easier to buy hardware than to look.
If the model is the business and you do not own the intelligence, you do not own your future. Every batch of proprietary data you hand to a foundation-model vendor might come back as the product that replaces you. The 3am one for a founder.
The whole industry's reflex, more GPUs, more data, more compute, is an expensive way to avoid the one unglamorous fix that actually works. The GPU bill is the symptom. The data is the disease. Nobody selling compute or scale will ever say it, which is exactly why it is the one to own.
Every training run on unrefined data is compute, money and months you never get back, on a model capped by its diet, at a moment when you cannot even get the GPUs. The waste is invisible until someone shows you the multiplier.
The friction points between what the builder carries and what the category keeps selling them.
DatologyAI is a data refinery for AI model builders. It does not sell data and it does not label it. It takes the data a company already owns and makes it dramatically better to train on, through an automated four-stage pipeline: clean, curate, create, compose. It runs in the customer's own cloud, so the data never leaves. The founder, Ari Morcos, is ex-Meta FAIR and DeepMind, and the company runs as a frontier data research lab that folds its published research straight into the product. The whole play is credibility: Ari, the research, and customer results. No hype, every number linked to a paper.
The output is the data to train a better model on far less compute. The proof is already on the table. Thomson Reuters mid-trained on a proprietary legal corpus and broke through the post-training ceiling on a budget under 1% of pre-training. Arcee built a frontier-class open-weights model with around ten people and 20 trillion curated tokens, competitive with models from 170-person teams that raised two billion dollars. That second number is the whole thesis in one line.
For five years the recipe was simple: scrape a slice of the web, clean it lightly, train on it. That era is closing. Research group Epoch AI projects that the stock of high-quality public text runs out around 2028, with the best language data depleting as early as 2026. Ilya Sutskever calls it "peak data" and compares it to fossil fuels. A Nature piece put it more bluntly: researchers have nearly sucked the internet dry. The Wall Street Journal's framing, that for data-guzzling AI the internet is simply too small, is now the consensus, not a hot take.
At the same time, compute is rationed. GPUs are scarce and expensive, and the reflex answer, buy more, is getting harder and less defensible as the market starts pricing efficiency over brute force. The DeepSeek moment made "more with less" respectable overnight. So the advantage is moving from who can scrape and spend the most to who can refine what they already own.
The buyer is a narrow, global, elite group: roughly a thousand companies worldwide training their own models. Two shapes. Startups where the model is the business, like Deepgram in voice or Hudson River Trading, who say plainly, "we need to own our intelligence." And startup-minded enterprises like Thomson Reuters and LexisNexis, fifty-year incumbents being bitten by AI natives, who have to move like a startup or become a museum. The category next door, data labelling, shows the stakes: when Scale AI sold a large stake to Meta, Google and OpenAI reportedly walked, unwilling to let a rival near their data. Owning your intelligence is not a slogan. It is already survival.
Ten techniques mapping what the category sounds like, looks like, measures, and misses.
Read twenty AI and data sites and the language sorts itself into the same six buckets. These are the phrases doing no work any more, the ones a sceptical, Twitter-native buyer has learned to scroll straight past.
The whole category measures itself in size. Bigness became the brag, which is exactly why bigness stopped being the differentiator.
Hype vocabulary that a technical audience actively distrusts. For this buyer, every one of these words lowers credibility, not raises it. Say the number, link the research, drop the adjective.
The oldest metaphor in the category, worn smooth. The refinery reframe is smart precisely because it takes the tired "oil" line somewhere true: you have drilled the field dry, so the value is in refining, not drilling.
The labelling world's badge of quality. It is also the admission of a ceiling: humans cannot read a trillion tokens. The proudest phrase in the category is a scaling limit dressed as a virtue.
Reassurance-by-badge. Necessary, but it describes the plumbing, not the payoff. It never once says what the buyer actually gets: a model that is theirs.
Everyone says "garbage in, garbage out" as a knowing shrug, then does nothing about it. The opening is to stop shrugging and make it the entire thesis. Weaponise the truism everyone repeats and no one acts on.
The opening for a credible challenger is to empty all six buckets and say the one thing none of them contain: it is not about how much you have, it is about what you do to it, and here is the research that proves it. In a category drunk on scale and hype, the sober, specific, proof-led voice is the loud one.
The category has one look. Glowing blue neural networks. Nodes and edges. A brain made of circuitry. Streams of particles flowing into the light. The abstract 3D orb or blob standing in for intelligence. Purple-to-blue gradients on everything. It signifies "AI" and communicates nothing, and it is indistinguishable from every other AI company, including the mystery spheres already sitting on the DatologyAI site.
Not one frame of it is about food, or refining, or the human who has to bet their company on this. The direction writes itself by opposition. No glowing brains, no abstract orbs, no gradient soup. Lean into the metaphors that are physical and true: the refinery, the kitchen, the diet, raw crude becoming jet fuel, real food versus junk. Show the builder, not the neural net. The category has never once made intelligence look like something you feed.
Tear down how this world announces itself and one film repeats: founder to camera, a benchmark chart, some lo-fi Silicon Valley b-roll, a synth sting, a Twitter thread. Scored out of ten for emotional pull.
The launches share one structure. The value is knowing it well enough to sit just outside it.
Ari on camera is right, the founder is the credibility. But if the film around him is the standard founder-plus-benchmark launch, it blends into the exact feed the founders are scrolling. The job is to keep the credibility and change the register: a launch that argues a point of view (models are what they eat) rather than announcing a product, and earns every claim with research. Witty, sober, proof-led. Not a benchmark chart with a face.
The whole category clusters in Safe and Rational: benchmarks, specs, pipelines. The rare warmth is a vague "democratise AI." The open quadrant is Bold and Human: a real argument about how models actually get good, carried with wit and a point of view, but still nailed to research. Datology's credibility gives it the unusual right to be bold there without tipping into hype.
Macro forces the launch can ride, resist or embody. Not trends, tensions.
The two humans who live with this: the founder who bets on the model, and the researcher who trains it. Composites drawn from AI founder threads, ML engineering forums, and the builder conversations behind this brief.
The founder talks about ownership, survival and being copied. The researcher talks about the grind and wanting to do the real work. Neither talks about parameter counts. The category markets in the language of scale. The buyer lives in the language of ownership and effort.
The category markets almost entirely to the researcher's scorecard: benchmarks, tokens, throughput. But the person who signs the cheque is measured on the founder's scorecard: a defensible moat and a business that survives. Datology wins the researcher scorecard on the research. The launch should be pitched at the founder's scorecard, ownership and survival, where nobody else is speaking.
Beyond the imagery, the meanings. What the category's signs are saying, and the archetype they all play.
The category plays the Magician: intelligence conjured from nothing, glowing and mysterious. The truer, unclaimed archetype is the Refiner, or the honest cook. Intelligence is not conjured. It is fed. Move from magic to nourishment, from the glowing orb to the thing on the plate, and you own a code the whole category has left on the table.
The category has two registers and nothing in between. One is dry academic benchmark-speak, loss curves and ablations, correct and cold. The other is breathless hype, the AGI-is-here register, which a technical audience distrusts on sight.
The buyer sits exactly in the gap: rigorous but human, sceptical but curious, allergic to hype but hungry for a point of view. They are on X watching founders argue and joke. Nobody is talking to that person with proof and personality at the same time.
Dry-witty and proof-led. Confident, a little irreverent, every claim earned by research. The smartest, least hyped voice in the room, the one that makes a sharp argument and then links the paper. It fits Datology's hard brand rule exactly: get attention without ever spending credibility.
Stop thinking like a marketer. Think like a therapist, a truth-teller, a troublemaker. Not summarising pain, provoking it.
A model is a compression of the data it was shown. Its ceiling is the information in that data. Parameters and compute only decide how fully it learns what is there; they cannot add what is not. So the data is the model's actual ceiling, and everything else, the architecture, the GPUs, the clever training tricks, is downstream of it. You cannot train past the quality of what you fed it. This is why "models are what they eat" is not a slogan. It is the physics of the thing.
Not a competitor. Not a model. The villain is a belief the whole industry has been high on for years: that intelligence can be bought by the pound. More data, more parameters, more GPUs, and quality will take care of itself. Call it the cult of more, the brute-force gospel, the scrape-it-all-and-sort-it-never era. It is the misreading of the bitter lesson that let everyone justify gluttony as strategy.
The cult of more is why estates of data sit unrefined, why teams buy GPUs instead of auditing inputs, why "quantity" got mistaken for "capability." It worked while the web was an open buffet. Now the buffet is closing, the compute is rationed, and the gospel is bankrupt. Naming that belief, and killing it, is the most useful and most ownable thing a credible challenger can do.
Surface: we need better training data.
Layer one: our model is not good enough and we do not fully know why.
Layer two: we cannot out-spend the giants on compute, so brute force is a losing game for us.
Layer three, the nerve: if we cannot build a model that is genuinely ours, we do not have a company, we have a wrapper on someone else's intelligence, and one day the market, the board and our own team will realise it. The real fear is not a weak benchmark. It is being exposed as thin.
The rational message: curated data raises signal per token, so you train a better model on less compute. Now wrap it in the frames a human actually moves on.
The category sells maximums: the biggest model, the most tokens, the highest benchmark. But the builder is not chasing a maximum. They are chasing enough.
The category has no native emotion, so borrow it from worlds where the same truth is felt in the body.
The diet. "Models are what they eat" is already Ari's line and the customer's instinct. It carries the whole thesis, quality of input decides quality of output, in an image anyone feels instantly, and it makes the GPU arms race look like buying a bigger plate.
"You know how a person is what they eat? An AI is exactly the same. It only knows what it was fed. Everyone's been feeding theirs the whole internet, junk and all, then buying bigger ovens to cook it faster. This company takes the food a business already has and makes it good enough that the AI actually gets smart, instead of just fat and expensive."
The test is the lean-in, and it comes at "buying bigger ovens to cook it faster." If a person outside tech smiles at that, the idea is clear enough to build on.
Under all of it: to own your intelligence, and to be freed from the drudgery so you can do the work you are proud of. The category sells throughput. Nobody sells ownership and relief, which is what the humans on both sides actually want.
For five years the recipe was to scrape a corner of the web, clean it lightly and train. The internet was free, vast and ungoverned, so the winning move was simply to grab more of it than the other team. Data was a thing you took, not a thing you made. That era built everything, and it is ending.
The good public web is picked clean and increasingly litigated. Compute stays scarce and expensive. The market starts rewarding return on compute over raw spend. In that world the moat shifts from who scraped the most to who refines best what they own, and from renting intelligence to owning it. The scrape era is closing. The refinery era is opening, and the companies that see it first will look, in hindsight, like they were early.
Ari Morcos built his career inside the scaling machine at Meta and DeepMind, and came out with a conviction he calls, in his own words, a bitter lesson: that the field's obsession with bigger models and more compute overlooks the thing that actually decides how good a model gets. Models are what they eat. Humans cannot curate at frontier scale, so it has to be automated, and done with research, not guesswork.
The deeper truth underneath the product is this. The frontier labs already win on data quality. They curate obsessively, and they simply do not sell it. So the game is quietly rigged: whoever has the best data pipeline wins, and only a handful of labs have one. Datology's whole reason to exist is to hand that same edge to everyone else, so the winner is not decided by who owns the most GPUs, but by who owns the best data.
Stop trying to buy intelligence by the pound. Own your data, refine it, and build the model only you can. Better data is the unfair advantage the big labs never wanted you to have.