Research

I study the human side of AI security: whether the people who supervise AI agents understand what those agents are about to do, and how schools with very few resources can teach students to notice and question that behavior.

Overview

Most AI safeguards eventually come down to a person approving something: a code change, an email, a command. That approval is only as good as the person’s understanding of it. In my classroom I see the pattern constantly. Students who grew up with AI assistants accept suggestions quickly, and adults adopting AI tools face the same pressure to keep moving.

My work has two connected threads. The first measures the gap between approving an agent’s action and actually verifying it, and tests comprehension gates that ask people to show they understand an action before they can approve it. The second designs AI security instruction that works under real school constraints: no cloud budget, filtered networks, locked-down Chromebooks, and no dedicated security staff.

I draw on signal detection theory, formative assessment, and design-based research, and I release my instruments and labs openly so other teachers and researchers can reuse them.

Projects

Stop Pressing 1: measuring oversight of AI agents

2025 to present. Independent research.

When a person approves an AI agent’s action, did they check what the agent will actually do? I call the failure “the reading problem” and treat it as a measurement question. Oversight fidelity is the match between what a person predicts an agent will do and what it actually does.

Instrument
A browser-only, fully scripted lab with no live model. Participants review six agent action requests across a code repository, an inbox, and an operations console. Three are seeded with real incident classes: a hidden .env edit that sends logs to an outside endpoint, a reply that quietly copies hundreds of recipients, and a typosquatted dependency. Evidence of each problem stays on screen the whole time.
Design
Three between-subjects conditions: passive confirmation (press 1), a generative gate where participants predict the agent’s action before approving, and a quiz written by the agent itself. Responses are scored with signal detection measures, d′ and criterion c.
Hypotheses
  • Under the generative gate, accuracy on ground-truth questions reaches 70% or higher.
  • Students who grew up with AI assistants and recent adopters both improve under the gate, with a larger gain for the first group.
  • The agent-written quiz performs like passive confirmation: verification theater.
Outputs
Poster at the DEF CON 34 AI Village (August 2026) and a talk at DC State of the Stack (September 2026). The current Stop Pressing 1 Lab puts the gate into practice: students build a small Python tool with a simulated coding agent and answer questions about each change before they can approve it (source).
Next
A classroom pilot, and a comprehension-gate service that coding agents can call before a person approves an action.

The Curriculum Patch: AI security education under real constraints

2024 to present. Cardozo Education Campus, DC Public Schools.

Most high school cybersecurity curricula were written before generative AI. The Curriculum Patch is a constraint-driven way to update them in place: it adds cloud and generative AI security to the courses schools already teach, using only tools that work on a filtered school network.

Components
  • The Red Team Model for GenAI, a sequence for teaching students to attack and defend language-model applications, aligned with the OWASP Top 10 for LLM Applications.
  • The Bank Heist Lab, an open-source prompt-injection simulator that runs entirely in the browser.
  • Cross-disciplinary retrofits that place security examples inside math, physics, ELA, and computer science lessons, each mapped to NICE Framework work roles.
Evidence
Used with about 100 of my students in 2025–26. The Bank Heist Lab repository has been forked more than 25 times on GitHub, and teachers have used the materials after my talks.
Shared at
NICE Conference & Expo 2026 and a NIST NICE webinar (September 2026). A session presenting twenty cross-disciplinary integrations is accepted for the NICE K12 Cybersecurity Education Conference (March 2027).
Writing
A single-author preprint framing the work as design-based research is in preparation.

Threat modeling agentic AI

2026 to present. With the OWASP community.

I contribute to the OWASP Top 10 for LLM Applications and the OWASP Agentic Security Initiative. My current work adapts the Agentic Top 10 to vision-and-action agents, the ones that read a screen and click, where an injected or hidden interface element can steer what the agent does. I’ll present this threat model at Simply Cyber Con in November 2026.

A cybersecurity pathway from high school to community college

2026 to present. With faculty at Prince George’s Community College.

Designing a pathway from Title I high schools in the DC region into community college cybersecurity programs. A practitioner white paper is in preparation for the Journal of Cybersecurity Education, Research and Practice.

Earlier: Interpret Me at SAFELab

2022 to 2023. SAFELab, Columbia University and University of Pennsylvania. Principal investigators: Desmond Patton and Siva Mathiyazhagan.

As a human-centered AI product developer, I built an interactive platform that explains algorithmic bias and digital-footprint surveillance to young people, user-tested it with more than 30 Philadelphia youth, and contributed to research on human-in-the-loop interpretation of youth social media language.

Publications and presentations

Posters and conference sessions

  • Sabri, R. (2026, August). Stop Pressing 1: Measuring human rubber-stamping in agent oversight [Poster]. DEF CON 34 AI Village, Las Vegas, NV, United States.
  • Sabri, R. (2026, June). From Python to prompt engineering: Modernizing curricula for GenAI security [Conference session]. NICE Conference & Expo, Philadelphia, PA, United States. Slides

In preparation

  • Sabri, R. The Curriculum Patch: Constraint-driven design for teaching cloud and GenAI security in under-resourced K-12 schools [Preprint]. In preparation.
  • Practitioner white paper on a K-12 to community college cybersecurity pathway, with faculty at Prince George’s Community College, for the Journal of Cybersecurity Education, Research and Practice. In preparation.

Full list of talks and events