Ask a general-purpose AI model to write a third-grade reading passage, and it will hand one back in seconds. Ask whether that passage actually sits at a third-grade reading level, matches the state academic standard a teacher is trying to teach, or whether the questions that follow measure the skill students are meant to practice, and the model’s tone won’t change. Even if it sounds sure of itself, it just might not be right, and students and teachers deserve more than a probabilistic result of what’s appropriate.

Better prompting won’t fix the core challenge: general-purpose AI models are trained to be broadly useful, not to know how a skill progresses from kindergarten through fifth grade, or which academic standard a given math problem is meant to teach. 

When AI gets it wrong, a teacher is usually the one who catches it, and likely loses confidence in the tool as a result. Teaching already asks a great deal of the people who do it, and depends on something no dataset can replicate: the critical relationship between a teacher and a student. The least AI can do is meet that work with the same rigor and evidence good teaching already runs on, instead of adding one more thing to double-check.

Our open tools and resources make proven teaching and learning practices easier to scale across the education ecosystem. Now, a year after our initial release of Knowledge Graph and Evaluators, we’re opening up a lot more for developers and educators.

Building AI for schools often surfaces three problems. Here’s what we’ve heard from developers, and what our products are built to resolve:

  • Giving an AI model relevant educational context.
  • Checking whether what it produces is actually good. 
  • Figuring out how to approach the task with the best teaching and learning practices. 

Below is what’s new on each of these three fronts.

1. Giving products reliable context for K-12 classrooms

Dataset catalog screen listing math, ELA, and social studies academic standards available for download
Knowledge Graph brings standards, learning components, progressions, durable skills, misconceptions, and other education data into one shared foundation.

Knowledge Graph gives developers structured educational data and context they can build from directly: what students are expected to learn, how skills build on one another, and where learning can break down. We’ve expanded that dataset foundation with refreshed academic standards across all 50 states, learning components across K–12 math and K–2 English language arts (ELA), and the new Eedi Misconception Graph, which adds more than 8,000 common math misconceptions informed by over 200 million real student answers.

In addition to core academic subjects, we’ve added new durable skills frameworks from the XQ Institute (XQ Competencies dataset) and the Carnegie Foundation for the Advancement of Teaching (Carnegie Skills Progression dataset), making skills like communication, collaboration, and critical thinking more concrete with detailed definitions, component skills, and developmental progressions available to developers.

Partners are already putting this context to work. 

  • Canva, a leading visual communication platform, uses Knowledge Graph datasets to build instructional content aligned to academic standards and the underlying learning components and progressions, helping teachers find resources that are instructionally sound and relevant to what they teach.
  • MagicSchool, a leading AI platform for schools, has worked with Learning Commons to explore how education-specific resources can support the development of high-quality AI tools for educators. The collaboration has helped inform MagicSchool’s thinking around educational standards and knowledge structures as it continues to build AI experiences designed specifically for K–12 schools.
  • OKO Labs, an AI-powered platform that facilitates small-group collaborative learning, is using Knowledge Graph to align its content to math learning components and to ground its feedback and measurement of durable skills like communication, collaboration, and critical thinking. Teachers and students get a research-backed read on both the math standards a group is working on, and the durable skills named in their district’s portrait of a graduate.
  • Really Great Reading (RGR), an education outcomes company changing the trajectory of literacy for students from kindergarten through high school, will release its next-generation Literacy Outcomes System (LitOS)™ this fall, which includes Knowledge Graph as a learning science–grounded component, ensuring teachers and students have the literacy skills needed to succeed.
  • TalkingPoints, which builds on the power of family-school partnerships to improve student outcomes, draws on Knowledge Graph curriculum and math learning components to generate lesson-aligned activities that connect what students are learning in class to academically grounded support at home, with instructions and activities translated into 150+ languages so every family can take part.

The common thread: instead of every developer, or every AI model, constructing educational knowledge from scratch, Knowledge Graph gives products a stronger, research-backed foundation of key educational datasets. That foundation will keep growing. We’ll continue to maintain, update, and expand these datasets with and for the field as academic standards and research evolve.

Karl Rectanus, CEO of RGR, sums up the need and urgency. “As we all focus on helping students learn in a rapidly modernizing world, shared public resources built on learning science like the Knowledge Graph help us move from pockets of excellence to systems of achievement. Our team at RGR is excited to truly partner with Learning Commons – both as a contributor and a beneficiary – and education leaders around the country on this important work to change the trajectory of literacy.”

2. Checking if output is ready for classroom use

General-purpose evaluation tools tell you whether a model is broadly performing well. They don’t tell you whether a reading passage is genuinely grade-appropriate, whether a math problem is aligned to the academic standard a teacher is trying to address, or whether a lesson activity is driving critical thinking. That’s a different, harder bar, and it’s the one our Evaluators are built to clear.

We build open-source, education-specific, LLM-as-judge evaluators with partner organizations like Student Achievement Partners (SAP), Achievement Network (ANet), Quill.org, and The Learning Agency to keep expert educators in the loop so that the evaluation rubric reflects a high benchmark for quality. We ask experts how they actually make a given judgment, turn that into a rubric, then test and refine the auto-evaluator against human annotation until experts agree it works as intended. 

Each evaluator assesses one dimension of quality rather than folding everything into a single score, because a single score hides exactly what a developer needs to know: which dimension is failing and what to fix. Our suite of Evaluators now cover grade-level appropriateness, text complexity, writing feedback, standards alignment, and durable skills to help developers measure whether activities or outputs meet an education-specific quality bar. 

  • The new Organizational Structure evaluator assesses how a body of text arranges its ideas and signals the relationships between them, so developers can tell whether a passage’s structure helps a reader follow the throughline or asks them to reconstruct it as part of the text’s design.
  • The Writing Feedback evaluator now measures additional dimensions of feedback to ensure it is specific, appropriate, and constructive. 
  • The new Critical Thinking evaluator surfaces where students are demonstrating skills to find, evaluate, and reason through evidence to build an argument, versus simply arriving at an answer.
  • The Math Standard Alignment evaluator determines whether content is aligned with the intended state academic standard to help ensure educators are using appropriate content for their intended lessons.

Our partners are already putting these evaluators to work. Level, a platform to help students build academic and life skills, is leveraging math and ELA learning components with Evaluators to ground learning activities and experiences in real academic expectations and instructional best practices. RightOn pairs educational context with the evaluation transparently to help raise teacher confidence in the AI-generated materials they use and give them the tools to make adjustments as needed. In both cases, Evaluators allow developers to test educational quality so teachers can trust the tools they use.

3. Guiding AI agents with expert judgment

Agent Skills provide expert-informed guidance for teaching and learning tasks, whether that’s helping an educator prepare a lesson, adapting it for a student who’s behind, or checking whether a class actually understood a concept before moving on. 

Agent Skills package research-backed instructional practices into open, reusable guidance you can build directly into an AI agent: a sequence of steps, things to watch for, and a quality bar for what a good result looks like. We build each one with subject-matter experts and iterate through repeated rounds of evaluation, so the guidance reflects how strong teaching actually happens, not just how it reads on paper. Our Agent Skills, co-developed with Anthropic include:

  • Lesson Plan Creation
  • Lesson Differentiation
  • Lesson Prep
  • Check for Understanding (refined with SAP

Later this year, we will release five new Agent Skills designed to support teacher planning. The new skills will range from helping audit lesson plans and formative assessments for quality, to helping support standards alignment adaptation. 

None of these skills prescribe a single teaching style or product experience, but instead provide a research-grounded starting point instead of a blank prompt. And every Agent Skill is open source and available on GitHub so you can inspect it, adapt it, or fork it. 

One open platform makes everything accessible

Dataset catalog screen listing all open Knowledge Graph datasets, including standards, learning components, and progressions
Explore Learning Commons tools in one place, from Knowledge Graph datasets and Evaluators to the access methods developers use to build with them.

The Learning Commons Platform brings all our tools together in one place, starting with a new Dataset Catalog that makes it easier to search, discover, and understand what’s available in the Knowledge Graph before you build with it.

Developers can use a free Learning Commons account to explore Knowledge Graph data, test content against our suite of Evaluators, and choose the various access methods that fit how they already work. Depending on the developer use case, that might mean downloading the full Knowledge Graph or a specific dataset, calling the Knowledge Graph via REST APIs or MCP, or integrating Evaluators via our SDKs.

Built for the field, with the field

We’re working with frontier model providers like OpenAI and Anthropic to bring educational grounding into general-purpose AI tools that are increasingly used in schools. We’re also announcing new partnerships with education organizations including Canva, Level, MagicSchool, OKO Labs, Really Great Reading, and TalkingPoints. 

While each of these organizations is working to address different classroom challenges, they’re able to leverage the same Learning Commons public infrastructure to support teaching and learning — from instructional content and literacy to family engagement, assessment, and teacher decision-making. 

Just getting started

As a nonprofit philanthropy, our mission is to make proven teaching and learning practices easier for educators and developers to build upon. We provide all our tools for free.

Today’s launch is a milestone, not an endpoint. A year of building alongside our partner organizations has sharpened what comes next for us. There’s still more educational context and data that AI needs to be more accurate and trustworthy. There are dimensions of pedagogy and rigor that we’re eager to work with education experts to shape how to measure quality in AI models. And there’s instructional knowledge sitting in published research and with expert educators that hasn’t yet been practically translated into public resources to improve the classroom tools teachers and students use every day. 

If you’re a researcher or an education organization, we’re eager to partner to bring more educational context and research into public infrastructure, to help raise the field as a whole. Together, we can collectively build a stronger education foundation and ecosystem where the tools educators and students rely on are worthy of the classroom and grounded in high-quality educational content, context, and proven teaching and learning practices.

This is what putting that mission into practice looks like. We’re just getting started.