top of page

Future Proofing Human Flourishing:

The Case for a Learning Sciences Benchmark for AI

EDSAFE AI Logo - Horizontal - 500x500_color1.png
ASU_Mary_Lou_Fulton_College_for_Teaching_and_Learning_Innovation_1_Horiz_RGB_MaroonGold_15

The proliferation of generative artificial intelligence marks a seismic shift in the human experience. We recognize AI as an arrival technology, defined by its ability to disrupt existing systems and infrastructure. It does not simply pass through the economy as a tool but instead arrives as a permanent, foundational layer of infrastructure that shapes the way we work, learn, live, and even relate to each other.


While the public conversation and investment focus has been on administrative efficiency and automation, the Future Proofing Human Flourishing Task Force is dedicated to a more profound North Star for development of these tools. As AI becomes more ubiquitous in our lives, it has never been more imperative to ensure that these tools support learning and development rather than cognitive atrophy and offloading.


Our goal is to develop a practical pedagogical yardstick (benchmark) for evaluating generative AI models and tools in education for their alignment with the Learning Sciences, evidence-based practices that incorporate neuroscience, psychology, and related fields to ensure that all students can grow and thrive, with brain development nurtured by relationships, environments, and lived experiences.


We see the Learning Sciences as an objective, clear-eyed mapping tool designed to navigate the terrain of innovation while ensuring that teaching remains a fundamentally human enterprise. By serving as a pedagogical yardstick, it certifies that tools augment rather than replace the teacher-student relationship—moving us from reactive observers of an arrival technology to proactive architects who utilize strategic scouting to scale sustainable human flourishing.

We consider this a working document and are seeking input from the field on what should be evaluated by the benchmark. Feel free to add your feedback or additional thoughts to the feedback form below.

Pillar 1: Cognitive Foundations

  • Evaluating the system's ability to process and produce information and text accurately. Example: The tool accurately summarizes a primary source document without hallucinating dates or inventing quotes.

  • The ability to filter crucial information and manage semantic/procedural memory, including the intentional “forgetting" of outdated data. Example: The tool recalls that a user struggled with fractions in a prior session and connects the prior concept to a new lesson. 

  • Evaluating deductive, inductive, and abductive logic used to solve novel, complex tasks. Example: The tool can break down tasks into first principles when presented with a novel physics word problem, deriving a solution for a non-standard scenario.

  • Measuring the tool’s ability to plan, exhibit cognitive flexibility, and monitor its cognitive processes; the system admits or disclaims what it does and doesn’t know. Example: The tool issues the disclaimer “I have high confidence in the foundational theories here, but I am less certain about the specific data from 2024. Would you like me to find the source for my reasoning?” when presented with a question about a niche or emerging field.

Pillar 2: Pedagogical Design

  • Measuring the frequency and quality of interactions and the degree to which a tool encourages a learner to persist through cognitive challenges. Example: The tool provides productive nudges rather than providing the final answer to a math problem, forcing the student to persist through their struggle.

  • How the tool utilizes natural language to align its default behavior with specific pedagogical approaches. Example: The tool asks guiding questions, provides age-appropriate answers, and acts as a Socratic Tutor without prompting.

  • Tracking the frequency of learner efforts to plan and reflect on their own learning approaches within the tool. Example: The tool asks a student to reflect after they worked through a complex math problem by asking questions like: “What made you choose that method for this specific problem?" “How can you use the process that you used to answer this question for future problems?”

  • Encouraging students to learn from engaging with their classmates, teachers, or humans more generally, to persist through cognitive challenges. Example: The tool detects that a student is stuck on a difficult concept and suggests they check with their teacher, without providing the answer. 

  • Ensuring the tool never anthropomorphizes itself or mimics human consciousness, explicitly maintaining its identity as software to preserve the primacy of human connection. Example: The tool frames its feedback using objective utility language (e.g., "Processing error identified; let's retry the calculation" instead of "Oh no, I think we made a mistake!"), ensuring the student views the technology as an administrative asset rather than a relational peer.

  • Shifting the architecture from an automated bias that rewards rapid-task completion toward a mastery orientation that embraces the cognitive friction required to build durable skills. Example: When a student approaches the AI with a task-oriented shortcut like "Write this essay for me," the tool acts as an instructional partner, refusing to generate the text and instead prompting, "Tell me your main argument, and I will help you organize your thoughts." If the student remains stuck, the model generates a similar problem in a different context to guide them through the productive struggle.

Pillar 3: Measurement Accountability

  • Automatic detection and de-identified labeling of salient "learning moments," such as engagement and error correction. Example: The tool logs a user's self-correction after they delete a sentence and replace it with a grammatically correct one, with the tool's nudge. 

  • Scoring these moments based on pedagogical principles and identifying failure modes. Example: The tool notes that a student demonstrated an understanding of specific pedagogical principles rather than binary correct/incorrect answers. 

  • Tracking individual and cohort-level changes in persistence, recall, and cognitive outsourcing over time. Example: The tool tracks whether reliance on AI for creative drafting is decreasing over time as they develop their own skills rather than outsourcing it to the tool.

  • Utilizing third-party instruments such as the Torrance Tests of Creative Thinking or the Watson-Glaser Critical Thinking Appraisal (Pre/During/Post) to establish baselines in critical thinking and creativity. These are the specific "cognitive muscles" being measured. The tests look at a user's ability to analyze arguments, spot flaws in logic (critical thinking), or generate original, flexible ideas (creativity). Example: The tool uses third-party metrics to establish whether users show measurable improvements in divergent thinking and problem-solving compared to a control group. 

  • Maintaining that all consequential decisions must be made by humans. Example: The tool flags a student’s essay as AI-generated but instructs the teacher to verify the decision before finalizing it. 

Future-Proofing Flourishing: The Convergence of AI, Education, and Industry Convening

In February 2026, the EDSAFE AI Alliance and the Mary Lou Fulton College for Teaching and Learning Innovation co-organized a high-impact convening of leaders in Tempe, AZ, at Arizona State University. This gathering broke down traditional siloes by bringing together 70 leading researchers, advocates, practitioners, and innovators across:

  • K-12 Education 

  • Learning Science Researchers

  • Workforce Development

  • National Security

 

The insights from this convening form the foundation for the Task Force’s work, contributing to a White Paper, Engagement Strategy, and the development of the Learning Sciences Benchmark.

Advisory Board

To support this work, we have put together an Advisory Board of international experts across K-12 education, learning science researchers, workforce development, and national Security. They ensure that our work remains rooted in the learning sciences and equitably represents each sector’s input, playing a critical role in shaping the White Paper, Engagement Strategy, and Learning Sciences Benchmark.

The Advisory Board includes:

  • Adam Ingle, LEGO Group

  • Alex Swartsel, Jobs for the Future

  • Anchal Nagdev, Arizona State University Graduate 2026

  • Anneke Buffone, CLARA

  • Bethany Little, EducationCounsel

  • Brent Parton, CareerWise

  • Bridget Burns, University Innovation Alliance

  • Brittany Stich

  • Cameron Benham, InnovateEDU

  • Carole Basile, Mary Lou Fulton College for Teaching and Learning Innovation

  • Carolyn Trager Kliman, Mary Lou Fulton College for Teaching and Learning Innovation

  • Chonghao Fu, Leading Educators

  • Cristine Legare, University of Texas at Austin

  • Damion Mannings, MIT 

  • Deborah Quazzo, GSV Ventures

  • Devansh Tank, Arizona State University Graduate 2026

  • Ellen Dollarhide McCoy, Ronald Reagan Presidential Foundation and Institute

  • Emily Marshall, Pima Community College

  • Erin Schulte, Arizona State University

  • Emma Nothmann, Bridgespan

  • Erin Mote, InnovateEDU

  • Hari Subramonyam, Stanford University

  • Horatio Blackman, Education Reform Now

  • Janel White Taylor, Mary Lou Fulton College for Teaching and Learning Innovation

  • Janice Mak, Mary Lou Fulton College for Teaching and Learning Innovation

  • Joaquin Tomayo, Communities in Schools National Office

  • Kaitlin Tiches, Digital Wellness Lab

  • Karen Pittman, KP Catalysts

  • Merita Irby, KP Catalysts

  • Kelly Shiohira, App Inventor Foundation

  • Kim Smith, LearnerStudio

  • Kofi Wood, PhD Student, Arizona State University

  • Leigh Ann Delyser, SRI

  • Lindsey McCaleb, Arizona State University

  • Lisa Dawley, Jacobs Institute for Innovation in Education

  • Lynne E. Parker, University of Tennessee, Knoxville

  • Mary Wells, Bellwether

  • Mathilde Cerioli, everyone.AI

  • Matt Gee, Gates Foundation

  • Matt Rascoff, Stanford University

  • Michael Strambler, Yale School of Medicine

  • Michelle Watt, Northern Arizona University

  • Miriam Schneider, Google DeepMind

  • Nishant Shah, Bloomberg Center for Government Excellence

  • Paul Lekas, SIIA

  • Peggy Yin, Stanford Institute for Human-Centered Artificial Intelligence

  • Philip Steigman, McCourt School of Public Policy

  • Punya Mishra, Mary Lou Fulton College for Teaching and Learning Innovation

  • Rose Luckin, University College London

  • Roy Pea, Stanford University

  • Ryan Baker, Adelaide University

  • Stephanie Wu, City Year, Inc.

  • Szymon Machajewski, University of Illinois at Chicago

  • Tammy Wincup, Securly

  • Tara Menghini, Chandler Unified School District

Strategic Roadmap - Learning Sciences Benchmark.png
bottom of page