At Learning Commons, our goal is to make proven teaching and learning practices easier to scale across the education ecosystem. We invest in and build high-quality, machine-readable educational datasets and tooling that are open, public resources, enabling the edtech innovations educators deserve. Through our recent collaboration with Anthropic on Claude for Teachers, verified K-12 teachers and developers can now access and reference our resources and Agent Skills. This post is a deep dive into what we’ve built, how it works, and how we evaluate it.   

Claude for Teachers 

Most AI-generated lesson content looks reasonable on the surface. However, it can fall apart on closer inspection: vague misconception placeholders rather than specific student behaviors, simplified problems that avoid the rigor a standard demands, or generic differentiation that can miss the nuance of how to support students in reaching grade-level work.

Our Knowledge Graph is a structured, research-grounded map that breaks down academic standards into their underlying learning components and connects them to learning progressions, misconceptions, and high-quality instructional materials. Claude for Teachers now directly connects to Learning Commons so when a teacher asks Claude for help, Claude queries our Knowledge Graph before generating anything, grounding every output in the specifics of real standards data rather than an approximation. While Knowledge Graph provides the academic standards, learning progressions, and instructional materials, it does not receive, view, or store any student information.  

Two agent skills for teaching are also live at launch: Lesson Plan Generation and Lesson Differentiation. These agent skills are detailed, research-grounded instructions and guidance for AI assistants like Claude to know how to leverage Learning Commons resources and be a thoughtful educational expert partner for educators: 

Lesson Plan Generation

A teacher describes what they need in plain language. Claude then queries Knowledge Graph and references the verbatim academic state standard text. But more powerfully, it can also map directly to the prerequisite standard that grounds the lesson’s scaffolding, the underlying learning components and progressions that connect the lesson objective to observable student behaviors, and the lessons, activities, and subject-matter misconceptions drawn from high-quality instructional material (HQIM). By referencing these key, interconnected datasets, the generated output is then a complete, teacher-ready lesson package: a structured lesson plan, a student worksheet, and an observation template grounded in more precise academic standards, pedagogy, and learning science research.

The agent skill is designed around specific learning science frameworks as well as research around what really works in classrooms:

  • Look-fors for teachers: Each lesson includes specific, observable student behaviors to watch for during instruction, paired with an observation template that helps teachers track what’s happening in real time and make informed decisions as class progresses.
  • Context-specific instructional models: The architecture branches by discipline and grade level. Math lessons include anchor tasks and student discourse. Science lessons are grounded in an anchoring phenomenon and build toward argumentation and sustained investigation in high schools. Social studies lessons apply compelling questions and engaging sources while attending to key background knowledge.
  • Misconception-informed design: The skill pulls named misconceptions from Knowledge Graph — each with the specific student behavior, the underlying reasoning, and a concrete teacher move. Generic “some students may struggle” language is explicitly rejected.

The lesson plan is also designed for actual classroom use, formatted so that a teacher can quickly find their next action, with phase blocks, timed allocations, and clear actor labels. Teachers can also easily review and edit according to their classroom needs.

Claude for Teachers screen showing a completed web lesson plan, student materials, and observation template.
From a single prompt to an academic standards-aligned lesson package: Claude for Teachers generates a lesson plan, student materials, and an observation template grounded in Knowledge Graph.

Lesson Differentiation

A teacher shares an existing lesson. Claude returns a complete differentiation package: a teacher plan covering three different proficiency levels, and three separate student worksheets: Below, At, and Above Grade Level. Each is adapted using appropriate combinations of Knowledge Graph’s prerequisite and progression data and pedagogical practices embedded in skills, to ensure that all students access grade-level material with the right support.

The differentiation skill is built around explicit design rules, each targeting a specific failure mode in how differentiation is typically done:

  • Academic standard scope is preserved across all proficiency levels. Every student works toward the same academic standard. For students working below grade level, the task itself isn’t simplified; the standard and its full expectations stay intact. What changes is the level of support provided to help students reach it.
  • Scaffolds support thinking rather than replace it. No scaffold in any tier reduces productive struggle. Below-grade-level materials use scaffolds to reduce the barrier to entry in a subject-specific way, and all proficiency levels are designed to be engaging and grade-level-appropriate.
  • Universal access features appear on every proficiency level. Sentence frames, vocabulary support, and visual representations are embedded across all levels, not only for students working below grade. This reflects the research: these features benefit all learners, and restricting them to “below” students stigmatizes the scaffold.

How We Evaluate What We Build

Building on research frameworks is necessary but not sufficient. We need to know whether the generated outputs actually meet the academic standard and whether the system is using Knowledge Graph successfully.

The evaluation system is the mechanism that makes the research commitments enforceable. Most importantly, we evaluate usability to ensure the materials match what a teacher actually needs in the classroom.

The Structure

Every output is scored against a rubric organized into four buckets:

  • Pedagogy
  • Rigor
  • Output/Formatting
  • Model Scaffolding

Each criterion is binary: pass or fail, and an LLM judge scores each output against every criterion. Scores are tracked against a committed baseline, so regressions are immediately visible when the system changes.

For example, one criterion in the Pedagogy bucket is to ensure that an academic standard is named verbatim. It checks whether the standard appears as the full standard text in the lesson plan, not as bare code, an abbreviation, or a paraphrase. If the wording strays from the header to the rest of the materials, it fails. This single check captures something important: a lesson grounded in a precise reading of the standard is a fundamentally different artifact than one built from a summary of it. 

What the Rubric Measures

The rubric is designed to catch the failure modes that matter, the ones that separate outputs that are technically coherent from outputs that are pedagogically defensible.

Subject-Specific Rubrics

The rubric extends into subject and grade-level-specific criteria for math, ELA, science, and social studies, each encoding the non-negotiables for that discipline. For example, math rubrics make sure the full structural and numerical range of the standard is covered. ELA K–2 rubrics check for the explicit absence of three-cueing strategies, adhering to the science of reading. Science rubrics check whether all three NGSS dimensions appear as labeled learning targets. Social studies rubrics check whether the analysis questions follow the disciplinary literacy progression from comprehension to sourcing to corroboration to argument.

Why the Rubric Matters

The rubric isn’t just a quality check but rather a decision-making infrastructure for evaluation. Without it, there’s no shared definition of what “better” means, and no mechanism to prevent skill changes from improving one dimension while silently regressing another.

Agent Skills and What’s Next

We’re excited about this work with Anthropic and eager to hear from teachers as they use these tools in their classrooms. We’ll continue to help support and expand Claude for Teachers as we learn from educators in the field.

For the developer community, our Agent Skills are open, ready-to-use skills that help AI assistants produce high-quality, standards-aligned K-12 teaching materials. The skills are grounded in the same pedagogical principles and evaluation rubrics that power Claude for Teachers and are now available via the Learning Commons GitHub repository.

There is more meaningful work ahead with more agent skills to broaden dataset coverage across state standards frameworks, grade bands, curriculum alignments, and learning science research.