The Real Problem Isn't Your Prompts
A department head at an international school showed me something last term that stuck with me.
She'd spent an evening generating AI-powered retrieval questions for her Year 10 biology unit. Forty questions. Clean formatting. Multiple difficulty levels.
She was proud of it.
Then a colleague checked the set before it went to students.
Seven questions contained factual errors. Two referenced a study that doesn't exist. One used terminology from a university-level course that would confuse her students.
"I spent two hours generating content," she told me. "Then three hours fixing it. I would have been faster writing from scratch."
She wasn't doing anything wrong. Her prompts were fine.
Her workflow was the problem.
No definition of what "excellent" meant before she started. No constraints on what the AI could invent. No verification step. No reuse system so the fixes would carry forward.
She had a tool. She didn't have a framework.
Why "Better Prompts" Won't Fix This
Every week, I see educators share prompt templates. "Use this for lesson plans." "Try this for rubrics." "Here's my magic prompt for differentiation."
And every week, the same failure modes repeat.
The reason is that two things are true about AI at the same time, and most workflows account for neither.
First, AI is an amplifier. Fei-Fei Li, co-director of the Stanford Institute for Human-Centered AI, calls it "the ultimate amplifier." It increases the volume and velocity of your existing practice. If your practice was mediocre, AI will produce mediocre work faster. If your curriculum design was weak before AI, AI will make it weaker, at scale. Deploying AI without a pedagogical foundation is a reliable way to fail.
Second, if the context is not provided, AI guesses — and it will never tell you it's guessing. NIST's generative AI risk profile names confabulation (what most people call hallucination) as one of twelve documented risk categories. Their definition: "the production of confidently stated but erroneous or false content."
That's not a rare glitch. It's a structural feature of how these systems work.
UNESCO's global guidance frames the challenge clearly: generative AI can be persuasive while wrong. Education needs human-centred capacity, safeguards, and teacher agency, not just faster output.
The implication isn't "stop using AI."
It's "stop treating AI like a vending machine and start treating it like a system that needs governance."
The only defence is to know what excellence looks like before you prompt, spec the work so precisely that AI has nothing left to guess, and evaluate every output against a standard you set — not one AI invented.
Professionals in every field spec their work before building: architects, engineers, designers. D.E.E.P. brings that discipline to learning design with AI.
The D.E.E.P. Framework (Bookmark This)
D.E.E.P. is not a prompting framework. It's an instructional design + risk + reuse operating system.
Each step has a human-led action, an AI-assisted action, and a reusable asset you keep.
| Step | What You Do | What AI Does | What You Keep |
|---|---|---|---|
| D: Define excellence | Name what students must do, what thinking must stay human, and what research says "good" looks like | Suggests objectives, misconceptions, exemplars | Objective + success criteria (Bloom-tagged), alignment table, risk tier, good/bad exemplars |
| E: Equip AI with rich context | Turn your definition into explicit, checkable rules and load every source AI needs | Drafts against the spec and self-critiques | Prompt template + rule set + context pack |
| E: Evaluate the output | Verify quality against your standard, then stress-test for every student | Generates a validation report and alternatives | TRUST report + equity audit + fixes log |
| P: Package the system | Standardise, share, set a review cycle | Helps package and document | Playbook / toolkit / template library |
Skip D and you can't evaluate — you're left guessing whether something "looks good." Skip the first E and AI fills every gap by guessing: confidently, plausibly, and wrong. Skip the second E and you ship work that's either weak or excludes your most vulnerable students. Skip P and the work dies with you.
This mirrors how serious risk management works. NIST's AI RMF Playbook organises actions around Govern, Map, Measure, Manage as ongoing functions, not a one-time checklist. D.E.E.P. applies the same logic to classroom-level AI use.
D: Define Excellence
Start here to avoid polished nonsense.
Most AI failures in education trace back to this: the teacher asked for content without defining what students need to think.
Before you open AI, become the expert on what excellence looks like — not just excellent output, but excellent learning.
Three questions that define excellence for this task
-
What should the student be able to DO after this? Not "know" — DO. A measurable capability. "Analyse the causes of the Cold War using at least two historiographical perspectives" is a learning goal. "Understand the Cold War" is a wish. If you cannot describe what the student does differently after the learning, you do not yet have a goal. You have a topic.
-
What thinking must remain human? Name the cognitive work that IS the learning. If the goal is "construct an argument," then AI constructing the argument removes the learning. If the goal is "evaluate sources," then AI evaluating the sources removes the learning. The thinking the student must do is the thinking AI must not do. This is the line between a scaffold and a wheelchair. A scaffold builds capacity. A wheelchair replaces it. Draw the line before you build anything.
-
Where does AI genuinely help? AI is excellent at generating raw material, levelling text for different readers, providing varied practice, giving first-pass feedback on structure, and producing multiple versions of the same thing. It is mediocre at making design decisions, judging quality, and understanding your specific students. It is actively harmful when it removes the productive struggle that builds understanding. If you cannot name where AI helps without crossing the line you just drew, you may not need AI for this task.
Use Bloom's to force cognitive precision
Instead of "create a lesson on photosynthesis," define:
"Students will analyse how light intensity changes the rate of photosynthesis using evidence from a graph."
Revised Bloom's taxonomy (Remember → Create + four knowledge types) makes objectives observable and assessable. That's exactly what AI needs to produce coherent learning materials, because vague objectives produce vague output.
Align outcomes, activities, and assessment
Constructive alignment means your outcomes, activities, and assessments point in the same direction. When they don't, students find workarounds — and AI makes those workarounds faster.
Fill this in before you generate anything:
- Outcome (Bloom verb + criteria):
- Activity that forces that thinking:
- Formative check (retrieval prompt / hinge question):
- Assessment evidence:
- Feedback plan:
Decide the risk tier before you generate
AI is not one level of risk. The risk changes with the task.
- Green: low-stakes drafting — lesson outline v1, examples, question stems. Always with teacher verification.
- Yellow: anything that can harm trust if wrong — assessments, feedback comments, parent communication, policy language. Requires the full TRUST pre-flight in the Evaluate step.
- Red: student PII, IEP/medical/discipline specifics, high-stakes decisions in open tools. AI should not touch this in most school contexts.
NIST highlights hallucinations, privacy, and information integrity as distinct generative AI risks. Decide before generation what tier you're in and what validation is required.
The European Commission's ethical guidelines for educators raise the same point: practical questions help educators with limited AI experience make safe decisions. Green/Yellow/Red is one way to operationalise that.
Research what excellence looks like in practice
Use research tools — Consensus, Perplexity Academic, Google Scholar, Elicit — to build a summary with four elements:
- What does research say excellence looks like for this task? If you are building a feedback tool, find what research says about effective feedback: timing, specificity, actionability. If you are building a lesson planner, find what the evidence says about how people actually learn, not what convention says lessons should look like.
- What does the evidence-based process look like, step by step? Not a vague description of best practice. A specific sequence a beginner could follow.
- What distinguishes excellent practice from weak practice? Collect real examples of both. Anti-examples define the boundaries AI must not cross: an objective that uses "understand" as its verb, feedback that praises effort without naming a strength, an assessment that tests recall when the goal is analysis.
- What are the common mistakes? The mistakes AI is most likely to make are the mistakes most people make — because AI learned from most people.
Two evidence anchors worth returning to:
Retrieval practice improves long-term retention. Roediger and Karpicke's 2006 study showed that students who practised retrieval significantly outperformed those who restudied — even though the restudying group felt more confident. AI can generate retrieval items quickly, but you must verify accuracy and schedule spacing.
Feedback is powerful but variable. A large meta-analysis found meaningful positive effects overall, but results ranged widely — including negative effects when feedback was poorly designed. AI can draft feedback stems. You must ensure they're specific, actionable, and aligned to criteria.
This step is non-negotiable. Without it you have no standard to evaluate AI's output against, and you'll be left judging whether something "looks good" — which is exactly how polished garbage gets approved.
The question to answer before moving on: "I know what the student should be able to do, what thinking must stay human, and what research says excellence looks like. I can recognise it when I see it — and I can recognise when it's missing."
E: Equip AI with Rich Context
Your prompt must function like a contract, not a wish.
Most educator prompts fail because they give AI freedom it shouldn't have. There are three jobs here.
1. Turn your definition of excellence into rules you can verify
Every finding from the D step should become a specific, testable instruction. "Make it high quality" is a wish. "Every learning objective must contain a measurable verb, a specific performance context, and a success criterion aligned to the IB assessment rubric" is a rule. If you cannot check whether a rule was followed by reading the output, rewrite the rule.
Rules fall into four categories:
- Content rules — what to include, what to exclude, what sources to draw from
- Format rules — structure, length, layout, style
- Behaviour rules — how to handle ambiguity, when to flag uncertainty, what to do when information is missing
- Constraint rules — boundaries not to cross, topics to avoid, assumptions not to make
Include rules that protect the line you drew in the D step. If the student's job is to construct the argument: "Do not write the argument. Provide three pieces of evidence the student could use, with source citations, and let the student construct the argument." If the student's job is to evaluate sources: "Present the sources without evaluation. Ask the student three questions that guide them toward judging reliability, perspective, and relevance."
The best test of a rule: could someone who has never spoken to you verify whether AI followed it? If not, the rule is too vague.
The four constraints that most improve reliability:
- Source boundary: "Use only the evidence and context below."
- Uncertainty behaviour: "If unknown or unsure, say so. Don't invent citations." (This directly addresses confabulation risk.)
- Output format: "Use the alignment table. Label any assumptions."
- Self-audit requirement: "After your output, critique it against the quality rules."
2. Structure your instructions so AI processes them clearly
Having excellent rules is not enough. How you deliver them affects what you get back. This is not about tricks — it's clear communication with a system that takes your instructions literally.
- Set the role precisely. Not "You are a helpful assistant," which produces generic output. "You are an experienced IB Diploma Programme Biology teacher designing formative assessments for a multilingual class of 25 students in years 12–13" gives AI a specific lens.
- Require reasoning before output. "Before producing your final output, think through your approach step by step. For each design decision, explain what options you considered, what evidence supports your choice, and what trade-offs you accepted. Show your reasoning BEFORE your final output, not after." When AI must show its work, it makes fewer arbitrary choices — and you can catch the ones it still makes.
- Specify the output format explicitly. If you need a table, describe the columns. If you need feedback comments, show one example of the format and tone you expect. Leaving format to AI produces whatever it defaults to.
- Tell AI what to do when it does not know. "If you lack sufficient information to answer accurately, say so. Do not guess. Flag any inference with [INFERENCE] so I can verify it."
- Use the examples and anti-examples from the D step. One strong example, annotated with what makes it strong, teaches AI more than a paragraph of abstract instruction.
3. Assemble everything AI needs
Upload or paste: the research summary, source documents (subject guides, rubrics, assessment criteria), learner context (who the students are, what they already know, what languages they speak, what challenges they face), format examples, anti-examples, and any prior work that should inform the output. Tell AI what each upload is and how to use it.
The core principle: if it is not in the prompt or uploads, AI fills the gap by guessing. It will not flag the gap. It will not ask for clarification. It will generate something confident, plausible, and possibly wrong — and move on. That is not a flaw in AI. That is how AI works. The only defence is to leave nothing to guess about.
E: Evaluate the Output
This is the difference between "AI use" and professional AI use.
Evaluation has two phases: check the quality, then check who it works for. You are not done until both are complete.
Phase 1: Evaluate against your standard
The first output is a draft, not a deliverable. It almost always contains at least one unsupported claim, one missed rule, or one piece of generic filler that sounds authoritative and says nothing.
Read it as a sceptic, not a collaborator. Then push back with specific prompts:
- "Check your output against each of my rules. For each rule, state whether you followed it fully, partially, or missed it — and show me the evidence."
- "What is the weakest claim in this output and why?"
- "What did you infer or assume that was not explicitly in the source material? Flag every inference."
- "What would a sceptical expert in this field challenge?"
- "Produce two alternative versions. Tell me which is stronger and explain why."
And critically, check against the line you drew:
- "Does this output do the student's thinking for them, or does it support the student in doing their own thinking?"
- "If a student used this, would THEY have done the cognitive work — or would the AI have done it?"
Budget for two to three rounds of structured push-back. Each round improves both the output and your spec. The refinements you make to rules during evaluation are often more valuable than the output itself — they make every future run better.
The TRUST pre-flight (7 checks, 3 minutes)
For anything in the Yellow tier, run this before it reaches students:
- T: Truth — factual accuracy. Definitions, names, steps. Are they correct?
- R: References — anything needing a source is flagged. No fabricated citations.
- U: Usefulness — concrete actions, not generic advice. Would a student know what to do?
- S: Standards — aligns to your Bloom level and constructive alignment table.
- T: Tone — age-appropriate, inclusive, safe for your context.
- Privacy — no student PII. Appropriate for your school's policies.
- Integrity — supports "thinking with AI," not outsourcing thinking to AI.
NIST explicitly names information integrity, confabulation, and privacy among the risks that must be managed. TRUST makes that management routine, not aspirational. For school-level privacy guidance, a pre-flight turns "responsible use" from a policy statement into a daily habit.
Phase 2: Stress-test for every student
Your output may be excellent by your standards and still fail your students. This phase catches that.
The Swap Test. Give your prompt or tool to a colleague. No verbal explanation. Can they use it and get a quality result? If they need you to explain what you meant, your spec is incomplete. Rewrite until the tool stands alone.
The Equity Audit. Test with three student profiles — not your average student, your edge cases:
- Your strongest, most experienced student. Does the tool still challenge them, or hand them a shortcut that removes the thinking they need?
- Your student with the most language barriers. Does the output assume native-level English? Does it use idioms, cultural references, or academic conventions that exclude multilingual learners?
- Your student with the most different cultural or educational background from the assumed default. Does the tool assume prior knowledge, family context, learning history, or values that do not apply?
If the tool works only for the middle of your classroom, it works for the students who needed it least.
The Bias Check. Read the output for five patterns:
- Language bias — does it default to Western academic English norms?
- Framing bias — does it present one perspective as neutral or default?
- Source bias — are cited examples and references from a narrow range of cultures?
- Omission bias — whose experience is missing entirely?
- Emotional bias — does it use language implying one outcome is obviously correct?
The Scaffold Check. Return to the line you drew in the D step. Does this build the student's ability to work independently over time, or create dependency? A scaffold you remove. A wheelchair you ride forever. If students cannot do the task WITHOUT the tool after using it, the tool is not teaching. It is replacing.
Keep a fixes log
Every time you catch an error, log it. After a month you'll have a clear picture of where AI fails in your subject, and you can feed those patterns into better rules (Equip) and better exemplars (Define).
The framework improves itself.
When to move on: "Would I stake my professional reputation on this — does it meet my standard AND does it work for every student in my classroom?"
If the answer is no, find where it fails. Weak output quality usually traces back to a thin definition of excellence (D) or vague rules (Equip). Output that excludes students means your rules don't account for the edge cases you just found. Go back and tighten.
P: Package the System
Once your inputs reliably produce excellent, equitable outputs, stop treating them as one-off conversations. Package the full spec — research, rules, context, examples, test results — into something your team can reuse.
Start with the simplest thing that works:
| Form | When to use it |
|---|---|
| Prompt template | You use it occasionally. Copy-paste into any AI tool. |
| Project with pre-loaded context | You use it weekly. Persistent workspace with documents already uploaded. |
| Custom GPT, Gem, or shared tool | Colleagues or students use it. They get consistent results without understanding the full spec. |
| App or agent | Simpler options create real friction. Not before. |
Build the simplest level first. Move up only when you hit genuine friction, not when it seems interesting to build something more complex. Over-engineering is the most common mistake at this step. A prompt template used every week is worth more than a custom GPT used once and abandoned.
Document what you built. Not for bureaucracy — for your future self and your colleagues. Record what research informed it, what rules you set, what edge cases you tested, what failed and what you fixed. When you revisit this in three months, that documentation is the difference between understanding your own work and starting over.
From individual practice to institutional capability
Individual teachers using D.E.E.P. get better results immediately. The real value multiplies when a department or school adopts it. Schools need consistent practice, not scattered heroics.
Use NIST AI RMF Playbook logic to keep governance continuous. The playbook frames actions around Govern, Map, Measure, Manage, and stresses that it isn't a one-size checklist — schools adopt what fits their context.
Use the TeachAI Toolkit to create guidance quickly. It includes principles and editable materials designed to help education systems develop responsible AI practices.
The urgency is real: RAND reports only 18% of U.S. principals said their school or district provided AI guidance, dropping to 13% in high-poverty schools. Most educators are operating without a system.
Packaging is what turns a few good individual practices into a school-wide playbook: shared prompt templates, verified exemplar banks, and a living fixes log that helps the whole team.
"I Don't Have Time for Four Steps"
You're right, if you treat them as four separate sessions.
But D.E.E.P. compresses. After the first run-through, most steps take minutes, not hours. The templates carry forward. The exemplar bank grows. The fixes log sharpens your rules automatically.
Here's the honest comparison:
Without D.E.E.P., you spend 45 minutes generating content, then 90 minutes fixing it — and none of those fixes carry forward to next week.
With D.E.E.P., you spend 30 minutes the first time (including verification), and 15 minutes each subsequent time, because the standard, the rules, and the context pack are already built.
The framework saves time. It just saves it on the second use, not the first.
Where D.E.E.P. Fits Among the Major AI Literacy Frameworks
Over the past two years, several major frameworks have defined what AI literacy looks like in K–12. But most of them are competency models. They describe what students and teachers should know and be able to do.
D.E.E.P. is different. It's an execution framework: how educators reliably design, validate, and scale AI-supported work.
That means they complement each other:
| Framework | What It Defines | What's Often Missing | How D.E.E.P. Complements It |
|---|---|---|---|
| Digital Promise (2024) | Understand / Evaluate / Use + justice-centred values | A rigorous verification + reuse workflow | D.E.E.P. supplies the production system: a defined standard, explicit rules, a two-phase evaluation, and packaging |
| UNESCO Student Framework (2024) | Student competencies across 4 dimensions with 3 progression levels | Translating competencies into weekly tasks and assessments | The Define step converts competencies into aligned learning experiences and feedback routines |
| UNESCO Teacher Framework (2024) | What teachers should master: pedagogy, ethics, AI foundations | "What do we do Monday?" repeatability | D.E.E.P. becomes the weekly operating routine for lesson design and safe AI use |
| U.S. DoE Toolkit (2024) | System leader playbook: risk, equity, policy, strategy (cites NIST RMF) | A simple instructional core workflow teachers actually reuse | D.E.E.P. becomes the instructional core mechanism inside the broader strategy |
| ETS Research Report (2025) | Learning progression with behavioural indicators (Emerging → Exemplary) | An implementation workflow without drift | D.E.E.P. provides the teacher workflow; ETS provides the progression spine |
| OECD/EC AILit (2025) | International competences across 4 domains (engage/create/manage/design) | Local execution habits and QA | Equip and Evaluate operationalise "manage AI" and "create with AI" through rules and trust testing |
The three gaps D.E.E.P. fills
1. A verification standard for hallucinations. Several frameworks emphasise evaluation and ethics. The U.S. DoE toolkit is explicit that hallucination risk must be managed. The Evaluate step is the operational muscle — it turns "evaluate AI" into a repeatable, auditable routine with an equity pass built in.
2. A learning-design alignment mechanism. Competency frameworks can end up taught as "AI content" rather than embedded across subjects. D.E.E.P. embeds AI literacy inside how lessons, tasks, and assessments are produced and validated.
3. A reuse-and-scale layer. TeachAI's toolkit is explicitly about building capacity, because most systems still lack it. The Package step turns a few good examples into a school-wide playbook and asset library.
The 30-Minute Weekly D.E.E.P. Sprint
This is the section you'll come back to.
One lesson. Four steps. Thirty minutes. Repeat weekly. Build a reusable library as you go.
- D (10 min): Tighten one objective with a Bloom verb + success criteria. Draw the line on what thinking stays human. Assign the risk tier. Complete the alignment table.
- E — Equip (8 min): Write the rules, set the role, add your uncertainty instruction, attach one example and one anti-example.
- E — Evaluate (9 min): Two rounds of push-back, then the TRUST pre-flight and one equity profile. Log any fixes.
- P (3 min): Save the prompt template and rules where your team can find them.
After four weeks, you have four verified, reusable lesson assets, and a growing fixes log that makes every subsequent sprint faster.
After a term, your department has a shared library that new staff can use from day one.
Staff PD Videos Worth 15 Minutes Each
If you're building team capability around D.E.E.P., these three talks pair well:
Sal Khan: How AI Could Save Education (TED) Use for: vision-setting. Why AI matters for learning.
UNESCO: Teacher to Teacher, AI Reshaping Education? Use for: equity and ethics. What we must protect — pairs with the Evaluate step's equity audit.
Ethan Mollick: Co-Intelligence Use for: practical habits. How to work with AI daily.
Watch one per staff meeting. Discuss for 10 minutes. Map the discussion to the relevant D.E.E.P. step.
Your Assignment
Pick one lesson you're teaching this week.
Run D.E.E.P. on it. All four steps. Thirty minutes.
Not as an experiment. As a test of the system.
At the end, you'll have a verified, aligned, reusable asset, and a clear sense of whether this framework works for your context.
If it does, run it again next week. And the week after.
The framework is designed to compound. Let it.
The Bottom Line
D.E.E.P. makes you better at knowing what learning needs, what AI should do, and what must remain human. The prompt is the delivery mechanism. The spec — the thinking behind it — is where the real work lives.
AI is the apprentice. You are the expert. The apprentice is fast, tireless, and confident. The apprentice also cannot tell the difference between a good lesson and a polished bad one.
That is your job. Act like it.
References
- UNESCO, Guidance for Generative AI in Education and Research (2023): unesdoc.unesco.org
- NIST, AI 600-1: Generative AI Profile (2024): doi.org/10.6028/NIST.AI.600-1
- NIST, AI RMF Playbook: airc.nist.gov
- UNESCO, AI Competency Framework for Teachers (2024): unesdoc.unesco.org
- UNESCO, AI Competency Framework for Students (2024): unesco.org
- Digital Promise, AI Literacy Framework (2024): digitalpromise.org
- U.S. DoE Office of EdTech, Empowering Education Leaders Toolkit (2024): eric.ed.gov
- ETS, AI Literacy Framework + Learning Progression (2025): rr.ets.org
- OECD/EC, AILit Framework Review Draft (2025): ailiteracyframework.org
- Biggs, Constructive Alignment: tru.ca
- Krathwohl, Revised Bloom's Taxonomy (2002): ihmc.us
- Roediger & Karpicke, The Power of Testing Memory (2006): wustl.edu
- Wisniewski et al., Power of Feedback Revisited: pmc.ncbi.nlm.nih.gov
- EC, Ethical Guidelines on AI for Educators: education.ec.europa.eu
- TeachAI Toolkit: teachai.org/toolkit
- RAND, AI Adoption Among Teachers and Principals: rand.org
- Student Privacy Compass, State AI Guidance: studentprivacycompass.org
