Not “it stopped working” — “each additional unit is buying you less than the last one did,” which is a much more specific and much more common problem.
The principle predates AI by centuries — it originates in classical economics, most associated with Turgot’s 18th-century writing on agricultural output and later formalized by Ricardo and Marshall: past a certain point, adding more of one input (labor on a fixed plot of land, in the original formulation) to a process with other inputs held fixed yields smaller and smaller additional output, not because the process breaks, but because the easy gains get captured first and what’s left gets harder to extract.
Applied to AI capability, “models have almost plateaued with slightly better efficiency than the last wave” is an informal diminishing-returns claim: the same increase in scale, data, or compute that produced a large capability jump in one generation produces a visibly smaller jump in the next. This is a distinct claim from Peak AI — diminishing returns is about the shape of the capability curve itself, while Peak AI is about whether investment and valuation have gotten ahead of that curve. A technology can show genuine diminishing technical returns while still being underpriced, or show strong continued technical gains while being wildly overvalued; the two claims aren’t the same and don’t rise or fall together.
Source (reference tier, not primary): “Diminishing returns,” Wikipedia, en.wikipedia.org/wiki/Diminishing_returns — cited as the standard modern restatement of the settled definition; the primary 18th–19th-century texts (Turgot, Ricardo, Marshall) predate DOI-era indexing and didn’t surface across three academic databases searched.