About Journify

Journify is an AI assistant built specifically for special education teams, designed to take paperwork off teachers’ plates and support planning so they can focus on their students. The platform automates progress tracking and goal generation and helps special education teachers create personalized instructional materials, including formative assessments, reading passages, writing supports, and more, aligned with each student’s individualized education program (IEP) goals and academic standards. Journify is used across 15+ states, supports 10,000+ students, and is certified by Digital Promise as both a Responsibly Designed AI Solution and an ESSA Research-Backed Product. Journify is also the winner of the K-12 track in the Tools Competition, which included 1,000+ companies, and the ASU+GSV Cup 50, which included 3,000+ companies.

Special education teachers bring deep pedagogical expertise to their work, but they’re asked to support students across a huge range of academic goals, often outside their content area of expertise. A teacher supporting a middle schooler with a literacy goal, for example, might have a clear sense of what that student needs, but with a full caseload of students across varied goals and grade levels, evaluating every AI-generated passage from AI tools takes time that could be spent with students. What that teacher needs is not another AI product that simply produces materials, but a tool that can generate appropriate texts with pedagogical nuance and quality, freeing the teacher to do what matters most: work with and support their students.

The problem: When AI generates instructional content, the burden of quality can’t fall on teachers alone

One of Journify’s core features is generating reading passages and literacy-based assessments calibrated to individual students. Through their work with teachers, the Journify team learned that a generated passage needed to be just right: neither too complex for a student to access nor too simple to be instructionally useful. Said otherwise, the texts Journify generated needed to be appropriate for the grade level of the students’ IEP goals.

This learning came early in the product’s life, as Journify gathered feedback directly from teachers to understand how generated content was landing. Teachers would flag passages that felt too long for a middle schooler or too simple for a high school student. It was a useful signal, but this assessment couldn’t easily scale. As Iain, Journify’s first engineering hire, put it, “You can’t improve what you can’t measure.”

Teachers are already asked to stretch across content areas while managing many competing demands. Asking them to audit AI-generated passages added yet another task to an already full plate. Journify recognized that if it was going to generate instructional materials for special education teachers, it also had a responsibility to systematically assess whether those materials were appropriate across grade levels and student contexts.

The solution: Build Evaluators directly into the development pipeline

When Iain came across the Learning Commons Evaluators, they recognized the measurement infrastructure their team was already looking for. The Grade Level Appropriateness (GLA) Evaluator measured the generated reading passages using a research-backed rubric to assess whether they were calibrated to the appropriate level for a given student. The Vocabulary Evaluator added another layer, measuring the accessibility and appropriateness of language within each generated passage. Having found the measurement infrastructure they were looking for, the Journify team got to work, implementing both evaluators into the tool they were already using to log and monitor AI-generated outputs. With lightweight, quick integration, every literacy generation that runs through the system is now automatically measured by these evaluators, with results tracked over time. More recently, the team added the Conventionality Evaluator, which scores text complexity along a spectrum from slightly complex to exceedingly complex.

Journify integrated all three evaluators into its existing traceability pipeline using Langfuse, a platform the team uses to log and monitor AI-generated outputs. Every literacy-related generation that comes through the system now runs automatically through each evaluator, with scores captured and tracked over time. The integration was lightweight enough to set up quickly and durable enough to run continuously at scale.

Iain describes what made the Evaluators immediately usable: “The learning science and the information inside the prompts are a goldmine, even independent of the scores. We like to measure things, but it’s also helpful to reverse-engineer them and back into the prompts we’re using to generate.” In other words, Evaluators did not just give Journify a way to score content. They gave the team a clearer language for what makes content pedagogically appropriate, and that understanding fed directly back into how they build their generation prompts.

Screenshot of a Journify AI-generated reading assessment on inference-making and citing text evidence.
Journify teachers review an AI-generated reading comprehension assessment before sending it to students. Evaluators run behind the scenes on the passage, giving teachers confidence that what they export is instructionally appropriate.

Early evidence: What measurement is making visible

The first thing the GLA Evaluator surfaced was a pattern that could now be quantified and tracked systematically. When Iain plotted grade-level appropriateness scores across K–12, the team identified meaningful variation across grade bands that had not previously been visible. The evaluator data made it possible to pinpoint where additional refinement would have the greatest impact. “That was the immediate value,” Iain said. “Now we know where to focus. We can say, let’s look at elementary school, let’s figure out what is going on.”

That kind of specificity changed how Journify works with its learning science consultant. Instead of bringing broad questions about content quality, the team could point to the exact grade bands in which its generated texts were less appropriate and ask targeted questions about why elementary passages might be drifting off level. Evaluators, as a suite of objective measurement tools, narrowed the search field and tightened the feedback loop between data and consultant expertise.

What’s next: Real-time measurement and action

Looking ahead, Journify plans to incorporate evaluator signals more directly into the generation workflow, helping ensure teachers receive high-quality materials with greater consistency at the moment of generation. The same measurement infrastructure that helps Journify’s team prioritize prompt improvements can also be leveraged to give teachers better information about each generated passage the moment it lands in their hands.

For a special education teacher deciding whether a generated reading passage is right for a student, the evaluator’s score can mean the difference between a quick check and a careful, line-by-line readthrough. That difference, multiplied across an entire caseload, is a real time saver that makes Evaluators useful for improving Journify’s internal development process. It also has a direct application at the moment of generation: giving teachers better information and greater confidence in what they receive.

Why it matters

Evaluators provide the research-backed quality criteria that warrant trust. By embedding them into Journify’s development pipeline, the team can hold its content generations to a consistent, evidence-informed standard rather than relying on anecdotal feedback. And as Journify moves toward surfacing evaluator signals directly to teachers, that same standard becomes a resource for teachers themselves, helping them understand not just what was generated, but whether it is ready for their students.

The goal of reducing paperwork and materials creation for special education teachers is only meaningful if the materials created are genuinely good. Journify’s embrace of measurement, using Evaluators, is a critical step in assuring teachers that what’s created for them is not only “good,” but also deserving of their students’ valuable instructional time.

Journify is a build partner of Learning Commons, integrating the Grade Level Appropriateness, Vocabulary, and Conventionality Evaluators into its AI generation pipeline to ensure the instructional materials it produces for special education teachers meet research-backed standards of quality.