Fall 2026 · CMSC 848N · University of Maryland

Generative
AI Agents

Foundations and frontiers of systems that reason, plan, act, use tools, and adapt in complex environments.

The course

Build agents whose claims can survive contact with the world.

An agent is more than a model. It is a governed sequential decision system: model, controller, environment, state and memory, tools and protocols, authority policy, evaluator, and people.

We will connect theoretical foundations to modern practice, with an emphasis on algorithmic building blocks, valid evaluation, and open research problems.

MeetingsTue & Thu · 11:00–12:15

CSI 3117 · America/New_York

PrerequisitesML + deep learning

Probability, optimization, linear algebra, and Python

CourseworkEvidence over demos

Diagnosis, implementation, and one personalized research project

Expected workload8–12 hours per week

Readings, homework, experiments, and team research

CommunicationELMS + course Slack

Course links and office hours will be shared through ELMS and course Slack

Schedule & slides

Thirteen modules.

Fall classes run August 31–December 11. CMSC 848N meets September 1–December 10; new lecture slide decks appear exactly two days before each 11:00 a.m. lecture. Dates follow the official UMD semester calendar.

No class Tuesday, October 13 (Fall Break). W7-L1 meets Thursday, October 15.

  1. Module01

    Agent-system foundations

    Define the complete agent system and reason about sequential experience.

  2. Module02

    Reinforcement learning foundations

    Complete the reinforcement-learning foundations needed for later agent training.

    L1
    Reinforcement Learning for Generative Agents (continued)Continued from prior lecture
    Pre-class reading →
    L2
    Reinforcement Learning for Generative Agents (continued)Continued from prior lecture
    Pre-class reading →
  3. Module03

    GRPO and alignment

    Connect group-relative reinforcement learning with preference optimization and controlled decoding.

  4. Module04

    Alignment continuation and planning

    Complete the alignment material, then study planning and state.

  5. Module05

    Agent harnesses, memory, and tool use

    Study how the harness selects context, maintains memory, and executes tool calls.

  6. Module06

    Durable composition

    Compose agents into workflows and justified multi-agent teams.

  7. Module07

    Inside a coding-agent harness

    Follow a real repair and explain how the harness shapes actions, feedback, and verification.

  8. Module08

    Training coding agents and improving their harnesses

    Learn from repair trajectories, then test changes to the runtime around a fixed model.

    Sources and reuse notices for the three coding-agent lectures
  9. Module09

    Research and scientific agents

    Connect sources, claims, hypotheses, experiments, and qualified conclusions.

    L1
    Deep Research Agents: Search, Evidence, and Reproducible SynthesisPDF release: Sun, Oct 25 · 11 AM
    Pre-class reading →
    L2
    Scientific Discovery Agents: Hypotheses, Experiments, and Valid ClaimsPDF release: Tue, Oct 27 · 11 AM
    Pre-class reading →
  10. Module10

    Embodied agents and robotics

    Design and evaluate closed-loop world-model and vision-language-action agents.

    L1
    World Models and Vision-Language-Action AgentsPDF release: Sun, Nov 1 · 11 AM
    Pre-class reading →
    L2
    Data-Efficient Embodied Learning and EvaluationPDF release: Tue, Nov 3 · 11 AM
    Pre-class reading →
  11. Module11

    Agent robustness and safety

    Threat-model the lifecycle and plan prevention, monitoring, and recovery.

    L1
    Adversarial Threats Across the Agent LifecyclePDF release: Sun, Nov 8 · 11 AM
    Pre-class reading →
    L2
    Defense in Depth, Monitoring, and Incident ResponsePDF release: Tue, Nov 10 · 11 AM
    Pre-class reading →
  12. Module12

    Learning from experience and self-improvement

    Govern persistent adaptation with evaluation, lineage, retention, and rollback.

    L1
    Learning from Experience: Prompts, Skills, Policies, and Continual AdaptationPDF release: Sun, Nov 15 · 11 AM
    Pre-class reading →
    L2
    Self-Modifying and Evolutionary AgentsPDF release: Tue, Nov 17 · 11 AM
    Pre-class reading →
  13. Module13

    Evaluation, provenance, and synthesis

    Make deployment decisions from reproducible evidence and accountable provenance.

    L1
    Agent Evaluation and Deployment EvidencePDF release: Sun, Nov 22 · 11 AM
    Pre-class reading →
    L2
    Provenance, Watermarking, and Frontier ChallengesPDF release: Sun, Nov 29 · 11 AM
    Pre-class reading →
  14. Upcoming—

    December 3 lecture

    The topic and reading assignment will be confirmed.

    Pending
    Topic to be confirmedTopic and readings pending
    Pre-class reading →
No classTue · Oct 13

UMD Fall Break

No classThu · Nov 26

UMD Thanksgiving Recess

Final meetingsDec 8 & 10 · 11:00 AM

Personalized research project presentations

Prepare for each lecture

Pre-class readings.

Start with Read first. Optional readings provide more detail or a different approach. Use the reading focus to guide your preparation.

These links are available ahead of the slide PDFs. Upcoming lecture readings are marked Planned and may be revised as the course develops.

Mathematical prerequisites · Lecture schedule

Updated . Titles identify the linked versions; arXiv years are initial posting years. Textbooks, engineering articles, and official specifications are labeled separately.

Module 01 · L2

Reinforcement Learning for Generative Agents

Reading focus: Write down an MDP and distinguish a policy, a value function, and a Bellman equation.

Read first

Module 02 · L1

Reinforcement Learning for Generative Agents (continued)

Reading focus: For the RL continuation, compare a sampled return with a bootstrapped value target.

Read first

Optional reading (1)
Module 02 · L2

Reinforcement Learning for Generative Agents (continued)

Reading focus: For the RL continuation, follow the policy-gradient derivation and explain the role of a baseline.

Read first

Optional reading (1)
Module 03 · L1

DeepSeek-R1, GRPO, and Agentic Reinforcement Learning

Reading focus: Compare PPO and GRPO: how are advantages estimated, and what does the reference policy do?

Read first

Optional reading (1)
Module 03 · L2

Alignment from DPO to Controlled Decoding

Reading focus: Derive the DPO loss from KL-regularized reward maximization, then compare parameter updates with control during decoding.

Read first

Optional reading (4)
Module 04 · L1

Alignment from DPO to Controlled Decoding (continued)

Reading focus: Continue the September 17 alignment material using the same slides and readings.

Read first

Optional reading (4)
Module 04 · L2

Planning and State

Reading focus: What state does a planner need, how does it predict an action's outcome, and when should it replan?

Read first

Optional reading (3)
Module 05 · L1

Memory and Context Engineering

Reading focus: Trace how the agent harness builds context for each model call. Then compare storing, retrieving, and summarizing experience. What evidence can each step lose?

Read first

Optional reading (12)
Module 05 · L2

Tool Use and Action Interfaces

Reading focus: How does an agent learn when to call a tool, and what must the tool interface make explicit?

Read first

Optional reading (6)
Module 06 · L1Planned

Agentic Workflows and Orchestration

Reading focus: Describe a workflow's branches, shared state, and recovery after a partially completed operation.

Read first

  • Workflow PatternsPaper · van der Aalst et al. · Distributed and Parallel Databases, 2003

    Focus on sequence, parallel split, synchronization, choice, and merge patterns. Publisher access may require a university login.

Optional reading (3)
Module 07 · L1Planned

Inside a Coding-Agent Harness

Reading focus: Follow a real repair from a failed reproduction through rejected edits to a submitted patch. How do the harness's tools, feedback, and stopping rules affect what the model can do and what the result verifies?

Read first

Optional reading (3)
Module 08 · L1Planned

Training Coding Agents

Reading focus: Turn executable repair tasks and interaction records into training data. How do trajectory selection, token loss masks, and verifier rewards determine what the model learns?

Read first

  • SWE-smith: Scaling Data for Software Engineering AgentsPaper · Yang et al. · arXiv 2025 · v2

    Follow task generation, validation, trajectory collection, and the reported training setup. Ask what each test and retained trajectory establishes.

  • SWE-smith: Train SWE-agentsDocumentation · SWE-smith project · official documentation

    Trace the steps from collected trajectories to supervised fine-tuning. Identify where successful-run filtering and training configuration enter.

Optional reading (4)
Module 08 · L2Planned

Optimizing Coding-Agent Harnesses

Reading focus: Hold the model fixed and propose a change to its runtime. What evidence shows that the change helps on new tasks after accounting for repeated selection, feedback, and compute?

Read first

Optional reading (4)
Module 09 · L1Planned

Deep Research Agents: Search, Evidence, and Reproducible Synthesis

Reading focus: How should a research agent gather sources, support its claims, and check citation quality?

Read first

Optional reading (2)
Module 09 · L2Planned

Scientific Discovery Agents: Hypotheses, Experiments, and Valid Claims

Reading focus: Distinguish proposing a hypothesis, running an experiment, and establishing a scientific result.

Read first

Optional reading (2)
Module 10 · L2Planned

Data-Efficient Embodied Learning and Evaluation

Reading focus: What transfers across robots and tasks, and what evidence supports transfer to a new setting?

Read first

Optional reading (2)
Module 11 · L2Planned

Defense in Depth, Monitoring, and Incident Response

Reading focus: What can a monitor observe, which actions can it stop, and what assumptions does the safety argument require?

Read first

Optional reading (2)
Module 12 · L1Planned

Learning from Experience: Prompts, Skills, Policies, and Continual Adaptation

Reading focus: What persists after feedback: a written reflection, a reusable skill, or a change to the model?

Read first

Optional reading (2)
Module 12 · L2Planned

Self-Modifying and Evolutionary Agents

Reading focus: How are proposed agent modifications generated, evaluated, and selected? How would you detect overfitting to the development tasks?

Read first

Optional reading (2)
Module 13 · L1Planned

Agent Evaluation and Deployment Evidence

Reading focus: What do repeated trials, changed conditions, and compute budgets reveal that one average success rate misses?

Read first

Optional reading (1)
Readings pending

Topic to be confirmed

The topic and reading assignment will be confirmed. No pre-class reading is assigned yet.

Personalized research project

Choose one project. Make it your own.

Professor Huang is designing a new catalogue of ambitious research candidates for the class. Each begins with a real open problem and a credible path toward a publishable contribution. Detailed briefs are available to enrolled students with the course password.

How project selection works

01

Review

Read the protected catalogue and identify several projects that fit your preparation and interests.

02

Rank

Rank three choices and propose one way to personalize your preferred direction. The planned deadline is Friday, September 11 at 5:00 p.m.

03

Form a team

Projects are completed in teams of up to three students. You may indicate preferred teammates when submitting your rankings.

04

Confirm

Final assignments will balance preparation, project demand, and available compute. Submission instructions will be posted in ELMS.

What each research project description includes

01

Motivation & gap

Why the problem matters and what unresolved gap makes it worth pursuing.

02

Question & novelty

A concrete research question and the seed of a new, defensible contribution.

03

Closest prior work

Starting references, relevant systems, and the work a new result must improve upon.

04

Research hypothesis

A technically credible first route, with space for the team to develop its own approach.

05

Data & baselines

Datasets, environments, comparison points, and practical starting assets.

06

Conference-level evidence

Metrics, budgets, strong baselines, ablations, and failure analyses for a valid claim.

07

Research artifacts

A conference-style paper, reproducible code, evidence, and a concise presentation.

Milestones

  1. Weeks 1–2Rank projects and form teams

    Submit project preferences, propose one personalization, and establish a shared workspace with a team of up to three students.

  2. Week 4Proposal

    Submit the task contract, benchmark or environment, and one-page plan.

  3. Week 6Baseline

    Deliver a runnable baseline and the first reproducible evidence bundle.

  4. Week 8Midterm report

    Report preliminary results and a concrete failure taxonomy.

  5. Week 10Freeze evaluation

    Lock held-out tasks, metrics, budgets, and contamination controls.

  6. Week 12Artifact audit

    Peer review, regression check, security test, and release rehearsal.

  7. Dec 8 & 10Present and submit

    Deliver the final presentation, paper-style report, runnable code, and supporting evidence.

Assessment

How your work will be assessed.

The assessment structure follows the Fall 2025 syllabus. Dates below are planned; the final Fall 2026 syllabus will confirm them, grade thresholds, and whether plus/minus grading is used.

10%

In-class quizzes

Impromptu short quizzes during lectures.

Sep 1–Dec 3
20%

Homework

Three assignments covering the course material.

Sep 30 · Oct 30 · Nov 30
30%

Midterm report

Project progress, a literature review, and baseline results.

Oct 15
20%

Final presentation

A concise presentation of the project, results, and limitations.

Dec 8 & 10
20%

Final report and artifacts

A research or engineering report, runnable code, and supporting evidence.

Dec 11

Project grading rubrics and submission requirements ↓

Project assessment

Project grading rubrics.

Proposed Fall 2026 rubrics, subject to confirmation in the final syllabus. Research and engineering tracks share expectations for technical depth, sound evaluation, and reproducibility.

Teams select a track and agree on scope with the instructor at the proposal milestone. Hybrid projects use one agreed rubric. The midterm package corresponds to the midterm report; final artifacts include the final report and supporting materials.

Each table totals 100 points within its component. The two final-artifact rubrics are alternatives. In-class quizzes (10%) and homework (20%) remain separate assessments.

Midterm package

30% of course grade

Midterm package: 100 points
CriterionPointsWhat earns full credit
Problem and scope20A precise question or use case, motivated by a concrete limitation; clearly defined success criteria and a feasible semester scope.
Related work and baseline25Accurate discussion of the closest approaches, verified citations, and a runnable baseline appropriate to the question.
Preliminary evidence30An end-to-end experiment or prototype, interpretable measurements, and an analysis of failures and uncertainty.
Evaluation plan and milestones25Held-out tasks, meaningful metrics, fair comparisons, resource budgets, major risks, and a credible plan for the remaining work.

Final presentation

20% of course grade

Final presentation: 100 points
CriterionPointsWhat earns full credit
Problem and approach20Clear motivation, necessary background, and an understandable explanation of the technical choices.
Evidence and interpretation40Results against appropriate baselines, readable figures, and conclusions supported by measurements. A demo supplements the evidence.
Limitations and discussion20Thoughtful treatment of failure cases, alternative explanations, uncertainty, and questions from the audience.
Communication and contributions20A coherent presentation within the allotted time, legible slides, and clearly identified contributions from each team member.

Final artifacts: research track

20% of course grade

Final artifacts: research track: 100 points
CriterionPointsWhat earns full credit
Research question and contribution20A precise hypothesis and a clear account of what the work adds relative to the closest prior work.
Experimental rigor35Strong baselines, controlled comparisons, relevant ablations, held-out evaluation, and uncertainty estimates where appropriate.
Technical correctness and analysis25Correct methods and mathematics, defensible conclusions, and substantive analysis of failures and competing explanations.
Reproducibility and reporting20A clear paper-style report, runnable code, documented configurations and dependencies, and sufficient details to reproduce the main results.

Final artifacts: engineering track

20% of course grade

Final artifacts: engineering track: 100 points
CriterionPointsWhat earns full credit
Correctness and reliability35An end-to-end system meeting agreed requirements, with automated tests, failure handling, and appropriate security and privacy protections.
Evaluation25Comparison with a simple baseline under documented workloads, measuring relevant quality, latency, cost, or usability outcomes.
System design20Well-justified interfaces and architectural choices, explicit tradeoffs, and a maintainable implementation.
Reproducibility and documentation20Reliable setup instructions, a repeatable demo, test commands, configurations, and clearly documented limitations.

What to submit

  • Midterm: a short report, runnable code, preliminary results, and an updated experimental plan.
  • Final presentation: slides explaining the problem, approach, results, limitations, and individual contributions.
  • Final artifacts: the research or engineering report, code, setup and evaluation commands, configurations, and supporting evidence. Engineering submissions also include automated tests and a repeatable demo.

Scoring and fairness

Score each criterion on a 0–4 scale, allowing half-points: 4 fully meets the criterion with convincing evidence; 3 largely meets it with minor gaps; 2 partially meets it with substantial gaps; 1 shows limited progress; 0 is missing or not assessable.

Criterion points = allocated points × score ÷ 4. For example, 80/100 on final artifacts contributes 16 of the 20 available course points.

  • Negative results can earn full credit when the investigation is rigorous and informative.
  • Conference acceptance, a novel algorithm, expensive compute, and positive results are not grading requirements.
  • Compare methods under fair resource budgets and explain unavoidable differences.
  • Disclose reused code, datasets, assets, AI assistance, and each member’s contributions.
  • Use synthetic or approved data for privacy-sensitive projects. Physical experiments require appropriate safety controls.
  • Agree on material scope changes before final evaluation.
  • Shared artifacts receive a team grade. Individual adjustments require documented contributions and demonstrated understanding, with an opportunity to clarify discrepancies.

Mathematical prerequisites

Review the foundations.

Use these reviews throughout the semester. For the papers assigned to individual lectures, see the pre-class reading list.

Week 1 · September 1 & 3

Refresh the mathematics behind this week’s lectures.

Review the topics you need before working through the reinforcement-learning definitions, Bellman equations, and policy-improvement derivations.

Course policies

Expectations and support.

These are planning guidelines. The official Fall 2026 syllabus will supersede this page where they differ. Students are responsible for reviewing UMD’s course-related policies and resources.

Teams & collaboration

Projects use teams of up to three students. Record individual contributions in the midterm and final reports. A self-proposed topic requires written instructor approval; individual quizzes and other designated individual work must be completed independently.

Generative AI use

AI tools may be used only when an assignment permits them. Disclose the model and version, preserve material prompts and tool traces, describe substantive edits, and verify every claim. AI assistance is not permitted on in-class quizzes unless explicitly stated.

Academic integrity

Follow the UMD Code of Academic Integrity and Honor Pledge. Cite external code, models, datasets, papers, prompts, and tools. Do not copy or distribute solutions, fabricate evidence, or present another person’s or system’s work as your own.

Late work

A shared 72-hour late bank applies across homework and written reports. It does not apply to quizzes, live presentations, or other in-person assessments. Once the bank is exhausted, late work is accepted only through an approved accommodation or excused-absence process.

Attendance & absences

Regular attendance and active participation are expected. Report known absences before the schedule-adjustment deadline and unexpected absences as soon as possible through a private ELMS or course-Slack message. Do not post medical documentation publicly.

Responsible experimentation

Use least privilege, sandbox side effects, protect private data, and obtain approval before tests that touch people or external systems. Human-subject research requires appropriate approval before data collection; consequential actions require monitoring and rollback ownership.

Accessibility & support

Students who need accommodations should contact the University’s accessibility service and the instructor early. Please raise barriers to course materials, assessments, or participation as soon as possible so arrangements can be made.

Communication & continuity

ELMS is the official source for announcements, assignments, grades, and urgent changes; course Slack supports discussion and help. If campus operations are interrupted, the continuation plan and any deadline changes will be announced through ELMS.

Planning page last updated October 6, 2026. Material changes after the first class will be dated and announced through ELMS.

CMSC 848N · Fall 2026

Models speak. Agents change state.

Our job is to understand the system in between—and demand evidence for what it can safely do.