We have done this before, and the name did real work
In the late 1990s somebody put four things together and called it LAMP. Linux, Apache, MySQL, and PHP or Perl or Python depending on who you asked. Nobody built a committee. It spread because it was useful, and it was useful because it compressed a long conversation into four letters.
Before LAMP, hiring someone to build a web application meant an hour of establishing what they had actually worked with. After LAMP, you said the word and both sides knew. The name did not create the stack. The stack already existed. The name made it possible to talk about, hire for, teach, and put on a CV.
MEAN did the same thing again around 2013 when the centre of gravity moved to JavaScript. MongoDB, Express, Angular, Node. Then React displaced Angular in enough places that MERN took over, and for most of a decade "MERN developer" was a genuinely informative thing to be.
Three names. Three eras. Each one made a labour market legible.
We are now several years into the largest change to knowledge work in a generation, and there is no name.
Why nothing has stuck
It is not for lack of trying. Various people have proposed AI stacks. A model plus a vector database plus a prompt layer. A particular combination of orchestration frameworks. Cloud vendors have their own versions, all of which are shaped suspiciously like the products that vendor sells.
None of them spread, and the reason is structural rather than a failure of marketing.
LAMP worked because the parts sat still. Apache was Apache for fifteen years. MySQL was MySQL. You could learn the stack and the stack would still be there when you finished learning it.
The AI tooling layer does not sit still. The vector database that was essential in 2023 turned out, for a large share of use cases, to be a Postgres extension. Orchestration frameworks have been adopted and abandoned faster than an acronym can spread. Any four-part name assembled from today's components describes a configuration that will be partly wrong within eighteen months, and everyone naming things knows it, which is why the names come out tentative and die quietly.
There is a second reason, and it is more interesting.
Standardisation is happening, but at the protocol layer rather than the component layer. Model context protocols for tool access. A de facto common shape for model endpoints. Emerging conventions for agents talking to other agents. That is a different kind of convergence, and it looks much more like how HTTP and REST won than like how LAMP won. Nobody ever said "the REST stack." REST was a shape everything took, not a set of parts you assembled.
So the component layer will not hold a name, and the protocol layer does not need one.
Which leaves the question that actually matters unanswered.
The category error
Here is the thing every attempt has got wrong, including the good ones.
They have all tried to name a system.
That made sense for LAMP, because in 1998 the only people who needed the word were the people building the system. The stack and the skillset were the same object seen from two angles. Naming one named the other.
That correspondence has broken.
A recruiter running candidates through an AI screening tool is not building a system. She is, however, doing four distinct things, all of which she can do well or badly, and none of which she has words for. She is deciding what information the tool gets to see. She is deciding which parts of her judgement to hand over and which to keep. She is forming a working theory of what the tool is reliable at. And she is deciding whether to trust the ranked list it hands back, knowing she will be the one explaining the hire in six months.
A radiographer does the same four things. So does a quantity surveyor, a claims adjuster, a paralegal, a marketing manager and a site engineer.
They are all doing AI-native work. None of them are developers. Almost none of them would use the phrase, and there is currently no way for any of them to say what they are good at, or for anyone to check.
LAMP described what was on the server. What we need now describes what is in the person.
That is the category error and correcting it is the whole move. Stop naming the system. Name the competence.
RAMP
Four layers. Retrieval, Agents, Models, Proof.
They are chosen on one criterion: they are the parts that do not churn. Tools underneath every one of them will be replaced repeatedly. The four things a person is doing will not.
R - Retrieval
Grounding the work in real context and data instead of plausible-sounding output.
The single most useful claim in this framework, and it is teachable in an afternoon:
Most AI failures are retrieval failures, not reasoning failures.
The model was not stupid. It could not see the thing it needed to see. It answered from general knowledge because the specific knowledge was not in front of it, and the answer came out fluent, confident, and untethered.
Someone strong at R can always tell you what the system had access to at the moment it answered. Someone weak at R cannot, and does not think of it as a question.
For an engineer this is retrieval architecture. What is indexed, how it is chunked, what is in scope at inference, what silently fell out of the context window.
For a financial analyst it is pulling the actual figures out of the actual system rather than accepting a generated summary of them.
For a lawyer it is the real clause from the real contract, cited, rather than a paraphrase that is right in tone and wrong in a way that matters.
For a marketer it is real customer language from real transcripts, instead of a persona the tool invented because nobody gave it anything better.
For a site engineer it is the current revision of the drawing rather than the one sitting in the shared folder.
Same competence. Wildly different surface. That pattern repeats for all four layers, which is the point.
A - Agents
Delegating and orchestrating work you do not do yourself.
Here is the observation that makes this land outside engineering, and it is the one I would lead with in any room of non-technical people:
This is management.
Decomposing a piece of work. Deciding which parts to hand over and which to keep. Putting a checkpoint where being wrong is expensive and leaving it out where it is not. Knowing what to do when the thing you delegated comes back confidently wrong.
Anybody who has run a team already holds most of this skill. They have simply never been told it transfers, and most of them assume the opposite, because the public story about AI is that it favours the young and the technical.
On the A layer specifically, it does not. It favours people who have spent years working out what to delegate to whom and where to put the check. That is a genuinely underrated asset and there are a lot of people sitting on it who currently think they are behind.
The failure mode is consistent across every job: delegating the part that carries the consequence and keeping the part that was easy. People hand over the judgement and keep the typing, because the typing feels like work. It is exactly backwards, and it is the most common mistake in the entire field.
M - Models
Knowing capability boundaries. What to trust, what to check, which tool for which job.
Not benchmark scores. Practical, domain-specific calibration: where does this thing confidently invent things in my field, and how would I recognise it?
Somebody strong at M can predict, before they run something, roughly how well it will go and where the failure will show up. That prediction is the skill. It is not knowledge of models in general, it is knowledge of how models behave against the specific texture of their own work.
A clinician who knows which kinds of summary are reliable and which need the source read has it. An underwriter who knows that the pattern matching helps on the familiar book and fails on a genuinely novel risk has it. An engineer choosing between a large model and a small one on cost and latency has it.
The failure mode is uniform trust. Somebody who treats every output with the same scepticism, whether that is high or low, has no calibration at all. They have an attitude toward AI rather than a working model of it. Both the enthusiast and the refuser fail this layer for the same reason.
P - Proof
Evaluating output and owning the outcome.
The scarce one. The one that matters. The reason this framework is worth anything.
Everyone teaches prompting. Prompting is a skill, it took about eighteen months to become common, and it will be fully absorbed into ordinary work within a few more. Almost nobody assesses whether a person can look at a piece of AI output in their own field and tell whether it is right.
That is the shortage. Two years of automating the production of work, and almost no investment in the capacity to check it.
P has a property the other three do not have, and it is the most important structural fact in this framework: P cannot be taught or assessed generically.
There is no such thing as being good at evaluating output in the abstract. You are good at evaluating financial models, or radiology reports, or Apex code under load, or a claims file, or a contract. The moment you leave the domain, the competence evaporates.
That is inconvenient for anyone selling a general course. It is also precisely what makes it worth certifying, because a domain-specific judgement is exactly the thing that cannot be faked, cannot be generated, and cannot be acquired in a weekend.
The failure mode: accepting output because it is well written. Fluency is not correctness. Fluency used to be weak evidence of competence, because producing fluent text took effort and understanding. Fluency is now free, and an enormous number of people have not updated on that. They are still reading polish as a signal, and it is costing them.
The same four layers, at different depths
The objection to a universal framework is always the same: if it applies to everyone it says nothing.
It survives that objection because the layers are constant and the depth is not.
An ML engineer working on R is building retrieval architecture, tuning chunk strategy, and measuring what got dropped. A claims adjuster working on R is making sure the model saw the full policy wording and not just the summary screen. These are not the same task and nobody would confuse them. They are the same competence: did the system have what it needed when it answered.
That is what makes RAMP a curriculum rather than a set of tracks. Everybody learns all four. The examples, the depth, and the assessment are drawn from the person's own domain.
Which, incidentally, is the only sane way to build a personalised course that still produces a comparable credential. Personalise the path. Standardise the bar. The route to the assessment is built for the individual. The assessment is not.
How you would actually test somebody
A framework that cannot be assessed is a poster. So here is the shape of it, in the order the layers should be tested.
R is tested by asking what the system saw. Give someone a piece of output and a scenario. Ask what it had access to, what it did not, and which part of the answer is therefore unreliable. People who have never thought about grounding cannot begin this question.
A is tested by decomposition. Give a real piece of work and ask what they would delegate, what they would keep, where they would put a checkpoint, and what they would do when it came back wrong. Answers cluster fast into people who have run work and people who have not.
M is tested by prediction. Before they run anything, ask where they expect it to fail. Then let them run it. Calibration is the gap between the prediction and what happened, and it is measurable.
P is tested by planting errors. Four outputs in their domain, three subtly wrong in ways that matter to a practitioner and are invisible to everyone else. Do they catch them, and can they say why?
That last one is the real assessment, and building it is real work. Every P item has to be written by somebody who knows the domain, and it only works for that domain. Fifty items for Salesforce data modelling do nothing for claims handling.
That cost is a feature. It is the reason a credential like this cannot be cloned by anyone in a weekend, and it is the reason it is worth something once built.
What this replaces
The old career ladder worked because titles were receipts. Senior Salesforce Developer was not a description of work, it was a claim that somebody watched you do the job for four years and signed off. That worked while jobs stayed still long enough for a title to keep its meaning.
Jobs do not stay still now. Roles get created, staffed, renamed and absorbed inside eighteen months. The receipt goes stale faster than you can earn it.
And the mechanism that used to produce the receipt is breaking at the bottom. The junior work that was the apprenticeship in almost every profession, the ticket queue, the document review, the audit sampling, the analyst model, the drafting, is exactly the work that compressed first. The ladder did not break at the top. It broke at the bottom, and it broke everywhere at once. Law, accounting, banking, software and design are all having versions of the same conversation without realising it is the same conversation.
So the question is not whether the ladder comes back. It is what replaces the receipt.
RAMP is a proposal for that. Four things anyone can be assessed on, in their own domain, at a stated depth, dated. Not a claim about how many years you sat somewhere. A claim about what you can actually do, in a form somebody else can check.
What could be wrong with this
A framework that does not publish its own weaknesses is marketing. Four honest ones.
The layers may not be equally weighted. The early evidence suggests P is worth more than the other three combined in most jobs. If that holds, RAMP is really a Proof framework with three supporting layers, and the honest version of this document in two years may say exactly that.
A and M may merge. As orchestration tooling matures, choosing a model may collapse into the orchestration layer and stop being a distinct human competence. If that happens the framework has three layers, not four, and it should be renamed rather than defended.
R may fade entirely. If context handling becomes genuinely automatic, retrieval could stop being a human skill within a few years, the way almost nobody manages memory allocation by hand anymore. That would be good for the world and bad for the acronym.
And it may simply not spread. Names do not stick by declaration. MERN spread because employers filtered for it, not because anyone campaigned. If RAMP does not end up in job specifications and delivery contracts, it will be a well-argued document that nobody uses, and the correct response to that is to drop it rather than push it.
This is a proposal, not a standard yet
Worth being explicit, because the space is full of frameworks asserting themselves into existence.
RAMP is a proposal. It is being published because the absence of any shared vocabulary is doing real damage to real people, most visibly to the ones outside technology who are experiencing all four of these layers with no words for any of them, no way to demonstrate competence, and no idea whether they are behind.
Four letters will not fix that. But a shared vocabulary is the precondition for everything else: hiring for it, teaching it, arguing about it, and proving it.
LAMP was not a standard either. It was four things somebody noticed kept appearing together, and it turned out to be useful enough that everyone adopted it.
The test for this one is the same. Not whether it is elegant. Whether a claims adjuster, a site engineer and a backend developer all read it and recognise their own job in it.
If you do, say so. If you do not, say what is missing. Both are more useful than agreement.
RAMP is maintained as an open framework by AiDOOS. RAMP Foundations, covering all four layers, is free and open to anyone. Assessment against a named domain is the Practitioner level and is what AiDOOS delivery pods are staffed against.