Fall 2026 · CMSC 848N · University of Maryland

Generative
AI Agents

Foundations and frontiers of systems that reason, plan, act, use tools, and adapt in complex environments.

The course

Build agents whose claims can survive contact with the world.

An agent is more than a model. It is a governed sequential decision system: model, controller, environment, state and memory, tools and protocols, authority policy, evaluator, and people.

We will connect theoretical foundations to modern practice, with an emphasis on algorithmic building blocks, valid evaluation, and open research problems.

InstructorProf. Furong Huang

IRB 4124 · furongh@umd.edu

MeetingsTue & Thu · 11:00–12:15

CSI 3117 · America/New_York

PrerequisitesML + deep learning

Probability, optimization, linear algebra, and Python

CourseworkEvidence over demos

Diagnosis, implementation, and one personalized research project

Expected workload8–12 hours per week

Readings, homework, experiments, and team research

CommunicationELMS + course Slack

Links, TAs, and office hours will be announced before classes begin

Schedule & slides

Thirteen modules.

Fall classes run August 31–December 11. CMSC 848N meets September 1–December 10; lecture slides appear exactly two days before each 11:00 a.m. lecture. Dates follow the official UMD semester calendar.

  1. Module01

    Agent-system foundations

    Define the complete agent system and reason about sequential experience.

    L1
    Agent Systems: The Model Is Not the AgentSlides Sun, Aug 30 · 11 AM
    L2
    RL Essentials for AgentsSlides Tue, Sep 1 · 11 AM
  2. Module02

    Learning agent policies

    Learn from demonstrations, preferences, and verifiable feedback.

    L1
    Learning Agent Policies from FeedbackSlides Sun, Sep 6 · 11 AM
    L2
    Training Multi-Turn Agents: Rollouts, Credit, and StabilitySlides Tue, Sep 8 · 11 AM
  3. Module03

    Reasoning and adaptive inference

    Control search, verification, compute, recovery, and stopping at test time.

    L1
    Reasoning at Test Time: Traces, Search, Verifiers, and ComputeSlides Sun, Sep 13 · 11 AM
    L2
    Adaptive Inference Control: Critics, Stopping, Recovery, and EfficiencySlides Tue, Sep 15 · 11 AM
  4. Module04

    Planning, state, memory, and context

    Build explicit state and memory that support controlled replanning.

    L1
    Planning and StateSlides Sun, Sep 20 · 11 AM
    L2
    Memory and Context EngineeringSlides Tue, Sep 22 · 11 AM
  5. Module05

    Tools, protocols, and authority

    Design typed tools with bounded authority, verification, and audit.

    L1
    Tool Interfaces and Action ProtocolsSlides Sun, Sep 27 · 11 AM
    L2
    Authorization and Safe ExecutionSlides Tue, Sep 29 · 11 AM
  6. Module06

    Durable composition

    Compose agents into workflows and justified multi-agent teams.

    L1
    Agentic Workflows and OrchestrationSlides Sun, Oct 4 · 11 AM
    L2
    Multi-Agent SystemsSlides Tue, Oct 6 · 11 AM
  7. Module07

    Web agents

    Build browser agents that ground actions and resist adversarial content.

    L1
    Web Agents: Perception, Control, and EvaluationSlides Tue, Oct 13 · 11 AM
    L2
    Securing Web Agents: Intent, Privilege, and Adversarial ContentSlides Sun, Oct 18 · 11 AM
  8. Module08

    Coding agents

    Navigate repositories, execute repair loops, and evaluate patches validly.

    L1
    Coding Agents: Repository Navigation and Repair LoopsSlides Tue, Oct 20 · 11 AM
    L2
    Coding Agents II: Execution, Training, and Valid EvaluationSlides Sun, Oct 25 · 11 AM
  9. Module09

    Research and scientific agents

    Connect sources, claims, hypotheses, experiments, and qualified conclusions.

    L1
    Deep Research Agents: Search, Evidence, and Reproducible SynthesisSlides Tue, Oct 27 · 11 AM
    L2
    Scientific Discovery Agents: Hypotheses, Experiments, and Valid ClaimsSlides Sun, Nov 1 · 11 AM
  10. Module10

    Embodied agents and robotics

    Design and evaluate closed-loop world-model and vision-language-action agents.

    L1
    World Models and Vision-Language-Action AgentsSlides Tue, Nov 3 · 11 AM
    L2
    Data-Efficient Embodied Learning and EvaluationSlides Sun, Nov 8 · 11 AM
  11. Module11

    Agent robustness and safety

    Threat-model the lifecycle and plan prevention, monitoring, and recovery.

    L1
    Adversarial Threats Across the Agent LifecycleSlides Tue, Nov 10 · 11 AM
    L2
    Defense in Depth, Monitoring, and Incident ResponseSlides Sun, Nov 15 · 11 AM
  12. Module12

    Learning from experience and self-improvement

    Govern persistent adaptation with evaluation, lineage, retention, and rollback.

    L1
    Learning from Experience: Prompts, Skills, Policies, and Continual AdaptationSlides Tue, Nov 17 · 11 AM
    L2
    Self-Modifying and Evolutionary AgentsSlides Sun, Nov 22 · 11 AM
  13. Module13

    Evaluation, provenance, and synthesis

    Make deployment decisions from reproducible evidence and accountable provenance.

    L1
    Agent Evaluation and Deployment EvidenceSlides Sun, Nov 29 · 11 AM
    L2
    Provenance, Watermarking, and Frontier ChallengesSlides Tue, Dec 1 · 11 AM
No classTue · Oct 13

UMD Fall Break

No classThu · Nov 26

UMD Thanksgiving Recess

Final meetingsDec 8 & 10 · 11:00 AM

Personalized research project presentations

Personalized research project

Choose one project. Make it your own.

Professor Huang is designing a new catalogue of ambitious research candidates for the class. Each begins with a real open problem and a credible path toward a publishable contribution. Detailed briefs are available to enrolled students with the course password.

How project selection works

01

Review

Read the protected catalogue and identify several projects that fit your preparation and interests.

02

Rank

Rank three choices and propose one way to personalize your preferred direction. The planned deadline is Friday, September 11 at 5:00 p.m.

03

Form a team

Projects are completed in teams of up to three students. You may indicate preferred teammates when submitting your rankings.

04

Confirm

Final assignments will balance preparation, project demand, and available compute. Submission instructions will be posted in ELMS.

What each research project description includes

01

Motivation & gap

Why the problem matters and what unresolved gap makes it worth pursuing.

02

Question & novelty

A concrete research question and the seed of a new, defensible contribution.

03

Closest prior work

Starting references, relevant systems, and the work a new result must improve upon.

04

Research hypothesis

A technically credible first route, with space for the team to develop its own approach.

05

Data & baselines

Datasets, environments, comparison points, and practical starting assets.

06

Conference-level evidence

Metrics, budgets, strong baselines, ablations, and failure analyses for a valid claim.

07

Research artifacts

A conference-style paper, reproducible code, evidence, and a concise presentation.

Milestones

  1. Weeks 1–2Rank projects and form teams

    Submit project preferences, propose one personalization, and establish a shared workspace with a team of up to three students.

  2. Week 4Proposal

    Submit the task contract, benchmark or environment, and one-page plan.

  3. Week 6Baseline

    Deliver a runnable baseline and the first reproducible evidence bundle.

  4. Week 8Midterm report

    Report preliminary results and a concrete failure taxonomy.

  5. Week 10Freeze evaluation

    Lock held-out tasks, metrics, budgets, and contamination controls.

  6. Week 12Artifact audit

    Peer review, regression check, security test, and release rehearsal.

  7. Dec 8 & 10Present and submit

    Deliver the final presentation, paper-style report, runnable code, and supporting evidence.

Assessment

How your work will be assessed.

The assessment structure follows the Fall 2025 syllabus. Dates below are planned; the final Fall 2026 syllabus will confirm them, grade thresholds, and whether plus/minus grading is used.

10%

In-class quizzes

Impromptu short quizzes during lectures.

Sep 1–Dec 3
20%

Homework

Three assignments covering the course material.

Sep 30 · Oct 30 · Nov 30
30%

Midterm report

Project progress, a literature review, and baseline results.

Oct 15
20%

Final presentation

A concise presentation of the project, results, and limitations.

Dec 8 & 10
20%

Final report

A paper-style account of the final project and its evidence.

Dec 11

Grading rubric

Reports

  • Problem statement clarity
  • Depth of literature review
  • Specific and credible research plan
  • Writing quality
  • Sound results and analysis in the final report

Grading rubric

Presentations

  • Time management
  • Clear and legible slides
  • Coherent oral communication
  • Novelty of the proposed idea
  • Comprehensive models, data, baselines, and ablations

Starting references

Read the primary work.

Core readings anchor the modules; further readings add context or alternative methods. Additional papers will be posted by week, with peer-reviewed work, preprints, official specifications, vendor evidence, and course synthesis clearly distinguished.

01Reinforcement Learning: An IntroductionCore · Module 01
Sutton & Barto · 2018
02A Survey on Large Language Model Based Autonomous AgentsFurther · Module 01
Wang et al. · 2024
03Proximal Policy Optimization AlgorithmsFurther · Modules 01–02
Schulman et al. · 2017
04Training Language Models to Follow Instructions with Human FeedbackCore · Module 02
Ouyang et al. · 2022
05Direct Preference OptimizationCore · Module 02
Rafailov et al. · 2023
06Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsFurther · Module 03
Wei et al. · 2022
07Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsFurther · Module 03
Yao et al. · 2023
08ReAct: Synergizing Reasoning and ActingCore · Modules 03 & 05
Yao et al. · 2023
09Generative Agents: Interactive Simulacra of Human BehaviorCore · Module 04
Park et al. · 2023
10ToolformerCore · Module 05
Schick et al. · 2023
11AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent ConversationCore · Module 06
Wu et al. · 2023
12WebArenaCore · Module 07
Zhou et al. · 2024
13AgentDojo: Evaluating Prompt Injection Attacks and Defenses for LLM AgentsCore · Modules 07 & 11
Debenedetti et al. · 2024
14SWE-benchFurther · Module 08
Jimenez et al. · 2024
15SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringCore · Module 08
Yang et al. · 2024
16The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoveryCore · Module 09
Lu et al. · 2024
17Mastering Diverse Domains through World ModelsFurther · Module 10
Hafner et al. · 2023
18OpenVLA: An Open-Source Vision-Language-Action ModelCore · Module 10
Kim et al. · 2024
19Reflexion: Language Agents with Verbal Reinforcement LearningCore · Module 12
Shinn et al. · 2023
20AgentBench: Evaluating LLMs as AgentsCore · Module 13
Liu et al. · 2023

Course policies

Expectations and support.

These are planning guidelines. The official Fall 2026 syllabus will supersede this page where they differ. Students are responsible for reviewing UMD’s course-related policies and resources.

Teams & collaboration

Projects use teams of up to three students. Record individual contributions in the midterm and final reports. A self-proposed topic requires written instructor approval; individual quizzes and other designated individual work must be completed independently.

Generative AI use

AI tools may be used only when an assignment permits them. Disclose the model and version, preserve material prompts and tool traces, describe substantive edits, and verify every claim. AI assistance is not permitted on in-class quizzes unless explicitly stated.

Academic integrity

Follow the UMD Code of Academic Integrity and Honor Pledge. Cite external code, models, datasets, papers, prompts, and tools. Do not copy or distribute solutions, fabricate evidence, or present another person’s or system’s work as your own.

Late work

A shared 72-hour late bank applies across homework and written reports. It does not apply to quizzes, live presentations, or other in-person assessments. Once the bank is exhausted, late work is accepted only through an approved accommodation or excused-absence process.

Attendance & absences

Regular attendance and active participation are expected. Report known absences before the schedule-adjustment deadline and unexpected absences as soon as possible through a private ELMS or course-Slack message. Do not post medical documentation publicly.

Responsible experimentation

Use least privilege, sandbox side effects, protect private data, and obtain approval before tests that touch people or external systems. Human-subject research requires appropriate approval before data collection; consequential actions require monitoring and rollback ownership.

Accessibility & support

Students who need accommodations should contact the University’s accessibility service and the instructor early. Please raise barriers to course materials, assessments, or participation as soon as possible so arrangements can be made.

Communication & continuity

ELMS is the official source for announcements, assignments, grades, and urgent changes; course Slack supports discussion and help. If campus operations are interrupted, the continuation plan and any deadline changes will be announced through ELMS.

Planning page last updated August 16, 2026. Material changes after the first class will be dated and announced through ELMS.

CMSC 848N · Fall 2026

Models speak. Agents change state.

Our job is to understand the system in between—and demand evidence for what it can safely do.