Who Educates the AI That Will Educate Our Students?

Who Educates the AI That Will Educate Our Students?

Educating AI is already an active practice — it is happening right now, just largely outside universities. For the past three years, the debate on AI in education has looked in only one direction: how much to let into the classroom. Ban or permit, detect or accept, resist or adapt. That gatekeeping debate matters, but it misses the larger question unfolding behind it.

The shift of the last two years explains why. Systems like ChatGPT or Claude no longer improve simply by ingesting more raw text. Their real evolution happens during post-training: the stage where models learn instruction-following, multi-step reasoning, error recognition, and specific behavioural profiles. Post-training is no longer a superficial alignment step. It is where a model’s core capabilities and pedagogical stance are forged.

To put it in educational terms: pre-training is general instruction, post-training is character formation. The former supplies content knowledge, the latter shapes dispositions. In teaching, dispositions are everything. An effective educator is rarely just someone with deep domain knowledge; they know when to hold back an answer, when to let a student struggle constructively, and how to tell productive friction from mere frustration. Post-training is precisely where these instincts are either built into a model or systematically erased.

Commercial AI models are deliberately trained to be efficient assistants: maximise speed, cut cognitive load, serve answers immediately. This runs counter to the intentional friction of good teaching, a loss I examined in an earlier article on the disappearance of the epistemic journey. Once an assistant is tuned to avoid friction at the training level, no amount of prompt engineering or classroom policy can put those instincts back.

Which brings us to the question this article is about. Educating AI is not a future possibility but a present activity — so who is shaping the educational values of the systems arriving in our schools?

Four Academic Approaches to Educating AI

Universities are already educating AI on their own terms, treating model training as an intentional educational project rather than a commercial product cycle. Four initiatives illustrate distinct academic priorities.

Open Methodology: Tülu

The Allen Institute for AI, working with the University of Washington, made the entire post-training pipeline for Tülu transparent — datasets, curation tools, decontamination scripts, code, evaluation suites. Crucially, the researchers also documented their failed experiments. Commercial labs keep dead ends quiet to protect competitive advantage; academic practice depends on reporting what does not work. Without that, hard-won insight into how models actually learn stays locked behind corporate doors.

Institutional Sovereignty: Apertus and Minerva

In Switzerland, EPFL, ETH Zurich and the CSCS developed Apertus, publishing model weights, datasets and fine-tuning methods alongside its explicit alignment principles. In commercial models, alignment criteria remain internal choices released as high-level summaries. Here they form a public, contestable record of values.

In Italy, the Sapienza NLP group led by Roberto Navigli took a similar curricular approach with Minerva, within the FAIR project. Rather than translating English datasets, the team authored native Italian instructions and mapped 21,000 training examples across a dedicated safety taxonomy. Structuring training material that deliberately is curriculum design in the fullest sense — treating the model as an ongoing research effort rather than a finished product.

Disciplinary Grounding: Meditron

EPFL and Yale built Meditron around disciplinary authority rather than raw scale. Its enduring contribution is the curriculum: roughly 47,000 clinical practice guidelines drawn from 17 recognised medical authorities, with 37,000 released openly from sources including the CDC, NICE and the WHO. Instead of web scraping, the model was trained on the literature the medical profession uses to train its own practitioners.

Learning Theory: ConvoLearn

Stanford’s Graduate School of Education built ConvoLearn to test how educational theory translates into training data. The dataset holds 2,134 tutor–student dialogues operationalising Scardamalia and Bereiter’s knowledge-building framework across six dimensions, among them metacognition, cognitive engagement and power dynamics. Authored by 323 experienced teachers, the resulting pedagogical signals correlated with expert evaluations of real classroom teaching. Proprietary initiatives such as Google’s LearnLM and TeachLM pursue pedagogy-informed training too, but their underlying data stays closed.

What Educating AI Actually Requires

These four projects share a principle that separates them from commercial frontier labs.

Commercial alignment is largely derived: objectives emerge from aggregated annotator preferences, benchmark tasks and usage metrics. What counts as a good response is settled by revealed preference, without a prior pedagogical framework.

Academic post-training starts instead from declared principles — established learning theories, clinical guidelines, published taxonomies, reproducible alignment standards. There is a source anyone can inspect, cite and contest.

In pedagogical terms, commercial training transmits an unexamined habitus; university-led training teaches an explicit curriculum. Only a curriculum can be evaluated, revised and held accountable by the community it serves.

Educating AI, then, is not a new burden placed on universities. It is the oldest thing they do, applied to a new kind of student — and it arrives precisely as the labour market is asking universities to justify what they are for.

Four Practical Tests for Any Post-Training Project

The distinction between derived and declared objectives is easy to state and harder to apply. What follows turns it into four questions you can put to any post-training project, academic or commercial, with an indication of where to find the answer.

They are independent. A project can pass three and fail the fourth, and how it fails usually tells you more than how many it passes.

1. Provenance: Where Does the Objective Come From?

Is there a citable document, written before the training run, stating what the model should be optimised toward?

This is not a question about open data. A dataset of annotator preferences records which option was clicked, not why. Declared objectives need a prior framework — a clinical protocol, a learning theory, a published safety taxonomy — that the training strategy can be held against.

Look for: a named framework in the methods section, or a separate document of principles. Absence is itself the finding.

2. Authorship: Are Domain Specialists Authors or Crowdworkers?

Does expert knowledge enter as an attributable position, or as an anonymised data point?

Three hundred credentialed teachers credited for designing instructional dialogues and three hundred anonymous annotators whose rankings collapse into a reward function are not the same methodological object, even when the judgements are identical. In education, pedagogical intent needs to be traceable to someone who can defend it.

Look for: author lists and acknowledgements. If experts appear only as a number, a qualification threshold or an agreement statistic, they were crowdworkers.

3. Falsifiability: Are Failed Approaches Documented?

Does the technical report include the configurations and ablations that produced nothing?

Commercial incentives discourage publishing failures; scholarly ones require it. The consequence is practical rather than moral. A method that fails in silence fails once for its authors, then again for everyone who tries it afterwards.

Look for: a section on unsuccessful approaches. Most technical reports do not have one.

4. External Validation: How Is Success Measured?

Is the system evaluated against a standard that existed beforehand, or against benchmarks built alongside the training data?

A benchmark built by the team that built the training set will tend to confirm what the training set encoded. This is not misconduct — it is circularity, and it is the default condition of the field. External validation is narrower and harder: an instrument developed for another purpose, measuring the construct rather than the artefact.

Look for: whether the evaluation instrument predates the project, and who built it.

Reading the Pattern

Applied to the four cases above, the criteria do not sort them cleanly, which is the point.

Meditron is strongest on provenance — clinical guidelines are as prior and citable as a source gets — while its evaluation stays largely inside domain benchmarks. Tülu sets the standard for falsifiability and is comparatively thin on provenance: its objective is a set of desirable capabilities, not a theory of what a model should be. ConvoLearn is the only one of the four to satisfy all four criteria, which is why it deserves an article of its own.

The criteria cut both ways. Google’s LearnLM passes on provenance, publishing an explicit learning science framework, and fails on authorship and falsifiability because its data and pipeline are closed. That is not a commercial project pretending to be academic. It is a genuinely mixed case, and a framework that could not represent it would not be worth much.

The point is not to sort public efforts from private ones. It is to identify what goes missing when alignment happens through implicit aggregation rather than open curriculum design — and therefore what an institution would have to supply to change it.

Three Counterarguments

Three objections deserve a hearing, and none of them settles the matter.

Performance. Academic post-training does not score better. A comparative evaluation of Italian-language models found that Italian-specific instruction tuning offered no measurable advantage, even on linguistic tasks, since base pre-training scale outweighed the fine-tuning data. Conceded — but standard benchmarks measure task execution against metrics the labs defined for themselves. None registers whether a model knows to withhold an answer. ConvoLearn’s correlation with real classroom observation suggests we have been testing the wrong thing.

Compute. Frontier training now demands tens of thousands of GPU hours for a single experiment, and public funding schemes tend to place universities behind an industrial lead. True, and beside the point. None of the four contributions above depended on matching compute. A fine-tuned checkpoint is obsolete within two years; a curated corpus or a validated instrument outlasts the hardware.

Structure. Misalignment may originate earlier, in pre-training, making post-training remedial. Grant it. Remediation is most of what education has ever been — no teacher meets students as blank slates. And pre-training corpora are built substantially on academic output, which means universities are already present at that stage as raw material and absent as designers. That describes the problem rather than excusing it.

The Stakes Are Not What Universities Fear

Universities are currently afraid of the wrong thing.

The fear is displacement — that models will teach, and institutions will become redundant. Understandable, and misdirected. Models will teach; that much is settled, and no admissions policy will change it. What remains open is whether anyone with a pedagogical mandate will have shaped what they teach, or whether that gets decided entirely by organisations answering to other pressures and publishing none of them.

The four criteria are diagnostic, but they double as a specification. Each names something an institution could supply, and none of it requires a frontier compute budget. Universities can publish normative frameworks — statements of what a model in their domain should be trained to do, written before any training run and meant to be cited and contested rather than to accompany a launch. They can curate disciplinary corpora, as medicine has already done, from the material a profession uses to form its own members. They can name experts as authors, so pedagogical intent stays traceable to someone who can defend it. And most urgently, they can build the measurement instruments that do not yet exist: validated ways to tell whether a system withholds an answer at the right moment, or leaves a learner able to work without it. Until those exist, no claim about pedagogical quality can be adjudicated — including the claims of the laboratories.

Which raises the question this article has avoided. Who, inside a university, does this?

The default answer is the computer science department, and the default answer is wrong. Departments with the technical capacity generally lack the theoretical grounding, and the reasoning turns circular: they build the benchmarks that measure what they already optimise for. Note where the strongest of the four cases came from — not an ML laboratory but a graduate school of education. Learning scientists, subject-matter faculty and assessment specialists hold the expertise post-training now demands, and have almost no institutional route into it. Building that route is an administrative problem rather than a research one, which is why nobody is doing it.

Universities can educate the systems that will educate their students. This is not a delegation of teaching but an extension of responsibility for it — and the responsibility does not stop at enrolment. These systems teach everyone: the professional checking a diagnosis, the parent looking something up, the citizen forming a view on a subject nobody ever taught them. Students at least have an institution around them that corrects, contextualises and assesses. Everyone else has the model alone. The most consequential teaching these systems do happens where no curriculum has ever reached.

The claim to that role is not opportunistic. It is genealogical. The mathematics of neural networks, the attention mechanism, the learning theory these systems implement — all of it was written inside universities, published through academic channels, and absorbed into the corpus the models were trained on. Universities are not encountering an external technology. They are encountering something they wrote.

There is a second reason, quieter and more practical. Curating a disciplinary corpus, coding instructional dialogue, building an instrument to measure teaching quality — none of this is work only faculty can do. It is exactly the kind of work through which students learn a discipline most deeply, because saying what counts as good practice is harder than performing it, and teaches more. A student who helps specify what a model should be trained to do has understood their field in a way no examination reaches. The loop closes properly here: universities form students who form the systems that go on to teach everyone else.

Institutions that have formed people since the Middle Ages are among the few actors equipped to state publicly, and then defend under criticism, what a system should be formed to do. That mandate is not shrinking. It has acquired a new kind of student — one that never graduates, and that begins teaching the world the moment its formation ends.

Cite this article: Cecchi, A. (2026). Who Educates the AI That Will Educate Our Students? Zenodo. https://doi.org/10.5281/zenodo.22012662

References

Alberto Cecchi

Educational Consulting, international academic relations. Researcher in the new media, design and social media. Specialties: 15 years of experience as Lecturer at the Universities of Urbino and Perugia, Multimedia Design and Computer Science. Author of Books and Articles about Design, New Media, Internet Security Education (History of Cryptography).

You may also like...