Instructors (TAs)
Fall 2026 · CMSC 848N · University of Maryland
Generative
AI Agents
Foundations and frontiers of systems that reason, plan, act, use tools, and adapt in complex environments.
The course
Build agents whose claims can survive contact with the world.
An agent is more than a model. It is a governed sequential decision system: model, controller, environment, state and memory, tools and protocols, authority policy, evaluator, and people.
We will connect theoretical foundations to modern practice, with an emphasis on algorithmic building blocks, valid evaluation, and open research problems.
CSI 3117 · America/New_York
Probability, optimization, linear algebra, and Python
Diagnosis, implementation, and one personalized research project
Readings, homework, experiments, and team research
Course links and office hours will be shared through ELMS and course Slack
Schedule & slides
Thirteen modules.
Fall classes run August 31–December 11. CMSC 848N meets September 1–December 10; new lecture slide decks appear exactly two days before each 11:00 a.m. lecture. Dates follow the official UMD semester calendar.
- Module01
Agent-system foundations
Define the complete agent system and reason about sequential experience.
- Module02
Reinforcement learning foundations
Complete the reinforcement-learning foundations needed for later agent training.
Reinforcement Learning for Generative Agents (continued)Continued from prior lectureReinforcement Learning for Generative Agents (continued)Continued from prior lecture - Module03
GRPO and alignment
Connect group-relative reinforcement learning with preference optimization and controlled decoding.
- Module04
Planning, state, memory, and context
Build explicit state and memory that support controlled replanning.
Planning and StateSlides Sun, Sep 20 · 11 AMMemory and Context EngineeringSlides Tue, Sep 22 · 11 AM - Module05
Tools, protocols, and authority
Design typed tools with bounded authority, verification, and audit.
Tool Interfaces and Action ProtocolsSlides Sun, Sep 27 · 11 AMAuthorization and Safe ExecutionSlides Tue, Sep 29 · 11 AM - Module06
Durable composition
Compose agents into workflows and justified multi-agent teams.
Agentic Workflows and OrchestrationSlides Sun, Oct 4 · 11 AMMulti-Agent SystemsSlides Tue, Oct 6 · 11 AM - Module07
Web agents
Build browser agents that ground actions and resist adversarial content.
Web Agents: Perception, Control, and EvaluationSlides Tue, Oct 13 · 11 AMSecuring Web Agents: Intent, Privilege, and Adversarial ContentSlides Sun, Oct 18 · 11 AM - Module08
Coding agents
Navigate repositories, execute repair loops, and evaluate patches validly.
Coding Agents: Repository Navigation and Repair LoopsSlides Tue, Oct 20 · 11 AMCoding Agents II: Execution, Training, and Valid EvaluationSlides Sun, Oct 25 · 11 AM - Module09
Research and scientific agents
Connect sources, claims, hypotheses, experiments, and qualified conclusions.
Deep Research Agents: Search, Evidence, and Reproducible SynthesisSlides Tue, Oct 27 · 11 AMScientific Discovery Agents: Hypotheses, Experiments, and Valid ClaimsSlides Sun, Nov 1 · 11 AM - Module10
Embodied agents and robotics
Design and evaluate closed-loop world-model and vision-language-action agents.
World Models and Vision-Language-Action AgentsSlides Tue, Nov 3 · 11 AMData-Efficient Embodied Learning and EvaluationSlides Sun, Nov 8 · 11 AM - Module11
Agent robustness and safety
Threat-model the lifecycle and plan prevention, monitoring, and recovery.
Adversarial Threats Across the Agent LifecycleSlides Tue, Nov 10 · 11 AMDefense in Depth, Monitoring, and Incident ResponseSlides Sun, Nov 15 · 11 AM - Module12
Learning from experience and self-improvement
Govern persistent adaptation with evaluation, lineage, retention, and rollback.
Learning from Experience: Prompts, Skills, Policies, and Continual AdaptationSlides Tue, Nov 17 · 11 AMSelf-Modifying and Evolutionary AgentsSlides Sun, Nov 22 · 11 AM - Module13
Evaluation, provenance, and synthesis
Make deployment decisions from reproducible evidence and accountable provenance.
Agent Evaluation and Deployment EvidenceSlides Sun, Nov 29 · 11 AMProvenance, Watermarking, and Frontier ChallengesSlides Tue, Dec 1 · 11 AM
UMD Fall Break
UMD Thanksgiving Recess
Personalized research project presentations
Personalized research project
Choose one project. Make it your own.
Professor Huang is designing a new catalogue of ambitious research candidates for the class. Each begins with a real open problem and a credible path toward a publishable contribution. Detailed briefs are available to enrolled students with the course password.
Protected project catalogue
Unlock the research projects.
The detailed catalogue is encrypted. Enter the password shared by Professor Huang to review the candidate questions, technical hypotheses, evidence requirements, and starting papers.
How project selection works
Review
Read the protected catalogue and identify several projects that fit your preparation and interests.
Rank
Rank three choices and propose one way to personalize your preferred direction. The planned deadline is Friday, September 11 at 5:00 p.m.
Form a team
Projects are completed in teams of up to three students. You may indicate preferred teammates when submitting your rankings.
Confirm
Final assignments will balance preparation, project demand, and available compute. Submission instructions will be posted in ELMS.
Fall 2026 candidate catalogue
Open a brief. Find the question you want to sharpen.
Rank several candidates and propose one personalization. Final scopes will reflect team preparation, available compute, and overlap across the class.
What each research project description includes
Motivation & gap
Why the problem matters and what unresolved gap makes it worth pursuing.
Question & novelty
A concrete research question and the seed of a new, defensible contribution.
Closest prior work
Starting references, relevant systems, and the work a new result must improve upon.
Research hypothesis
A technically credible first route, with space for the team to develop its own approach.
Data & baselines
Datasets, environments, comparison points, and practical starting assets.
Conference-level evidence
Metrics, budgets, strong baselines, ablations, and failure analyses for a valid claim.
Research artifacts
A conference-style paper, reproducible code, evidence, and a concise presentation.
Milestones
- Weeks 1–2Rank projects and form teams
Submit project preferences, propose one personalization, and establish a shared workspace with a team of up to three students.
- Week 4Proposal
Submit the task contract, benchmark or environment, and one-page plan.
- Week 6Baseline
Deliver a runnable baseline and the first reproducible evidence bundle.
- Week 8Midterm report
Report preliminary results and a concrete failure taxonomy.
- Week 10Freeze evaluation
Lock held-out tasks, metrics, budgets, and contamination controls.
- Week 12Artifact audit
Peer review, regression check, security test, and release rehearsal.
- Dec 8 & 10Present and submit
Deliver the final presentation, paper-style report, runnable code, and supporting evidence.
Assessment
How your work will be assessed.
The assessment structure follows the Fall 2025 syllabus. Dates below are planned; the final Fall 2026 syllabus will confirm them, grade thresholds, and whether plus/minus grading is used.
In-class quizzes
Impromptu short quizzes during lectures.
Sep 1–Dec 3Homework
Three assignments covering the course material.
Sep 30 · Oct 30 · Nov 30Midterm report
Project progress, a literature review, and baseline results.
Oct 15Final presentation
A concise presentation of the project, results, and limitations.
Dec 8 & 10Final report and artifacts
A research or engineering report, runnable code, and supporting evidence.
Dec 11Project assessment
Project grading rubrics.
Proposed Fall 2026 rubrics, subject to confirmation in the final syllabus. Research and engineering tracks share expectations for technical depth, sound evaluation, and reproducibility.
Teams select a track and agree on scope with the instructor at the proposal milestone. Hybrid projects use one agreed rubric. The midterm package corresponds to the midterm report; final artifacts include the final report and supporting materials.
Each table totals 100 points within its component. The two final-artifact rubrics are alternatives. In-class quizzes (10%) and homework (20%) remain separate assessments.
Midterm package
30% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Problem and scope | 20 | A precise question or use case, motivated by a concrete limitation; clearly defined success criteria and a feasible semester scope. |
| Related work and baseline | 25 | Accurate discussion of the closest approaches, verified citations, and a runnable baseline appropriate to the question. |
| Preliminary evidence | 30 | An end-to-end experiment or prototype, interpretable measurements, and an analysis of failures and uncertainty. |
| Evaluation plan and milestones | 25 | Held-out tasks, meaningful metrics, fair comparisons, resource budgets, major risks, and a credible plan for the remaining work. |
Final presentation
20% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Problem and approach | 20 | Clear motivation, necessary background, and an understandable explanation of the technical choices. |
| Evidence and interpretation | 40 | Results against appropriate baselines, readable figures, and conclusions supported by measurements. A demo supplements the evidence. |
| Limitations and discussion | 20 | Thoughtful treatment of failure cases, alternative explanations, uncertainty, and questions from the audience. |
| Communication and contributions | 20 | A coherent presentation within the allotted time, legible slides, and clearly identified contributions from each team member. |
Final artifacts: research track
20% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Research question and contribution | 20 | A precise hypothesis and a clear account of what the work adds relative to the closest prior work. |
| Experimental rigor | 35 | Strong baselines, controlled comparisons, relevant ablations, held-out evaluation, and uncertainty estimates where appropriate. |
| Technical correctness and analysis | 25 | Correct methods and mathematics, defensible conclusions, and substantive analysis of failures and competing explanations. |
| Reproducibility and reporting | 20 | A clear paper-style report, runnable code, documented configurations and dependencies, and sufficient details to reproduce the main results. |
Final artifacts: engineering track
20% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Correctness and reliability | 35 | An end-to-end system meeting agreed requirements, with automated tests, failure handling, and appropriate security and privacy protections. |
| Evaluation | 25 | Comparison with a simple baseline under documented workloads, measuring relevant quality, latency, cost, or usability outcomes. |
| System design | 20 | Well-justified interfaces and architectural choices, explicit tradeoffs, and a maintainable implementation. |
| Reproducibility and documentation | 20 | Reliable setup instructions, a repeatable demo, test commands, configurations, and clearly documented limitations. |
What to submit
- Midterm: a short report, runnable code, preliminary results, and an updated experimental plan.
- Final presentation: slides explaining the problem, approach, results, limitations, and individual contributions.
- Final artifacts: the research or engineering report, code, setup and evaluation commands, configurations, and supporting evidence. Engineering submissions also include automated tests and a repeatable demo.
Scoring and fairness
Score each criterion on a 0–4 scale, allowing half-points: 4 fully meets the criterion with convincing evidence; 3 largely meets it with minor gaps; 2 partially meets it with substantial gaps; 1 shows limited progress; 0 is missing or not assessable.
Criterion points = allocated points × score ÷ 4. For example, 80/100 on final artifacts contributes 16 of the 20 available course points.
- Negative results can earn full credit when the investigation is rigorous and informative.
- Conference acceptance, a novel algorithm, expensive compute, and positive results are not grading requirements.
- Compare methods under fair resource budgets and explain unavoidable differences.
- Disclose reused code, datasets, assets, AI assistance, and each member’s contributions.
- Use synthetic or approved data for privacy-sensitive projects. Physical experiments require appropriate safety controls.
- Agree on material scope changes before final evaluation.
- Shared artifacts receive a team grade. Individual adjustments require documented contributions and demonstrated understanding, with an opportunity to clarify discrepancies.
Starting references
Read the primary work.
Begin with the Week 1 mathematical prerequisites below. Core readings anchor the later modules; further readings add context or alternative methods.
Week 1 · September 1 & 3
Refresh the mathematics behind this week’s lectures.
Review the topics you need before working through the reinforcement-learning definitions, Bellman equations, and policy-improvement derivations.
PDF↗02Linear Algebra ReviewPrerequisite · Calculus & linear algebra
PDF↗03Advanced Linear Algebra ReviewPrerequisite · Calculus & linear algebra
PDF↗04Convex Analysis ReviewPrerequisite · Optimization
PDF↗05Optimization ReviewPrerequisite · Optimization
PDF↗06Probability ReviewPrerequisite · Probability & statistics
PDF↗
Across the semester
Course references.
Additional papers will be posted by week, with peer-reviewed work, preprints, official specifications, vendor evidence, and course synthesis clearly distinguished.
Reference metadata checked September 8, 2026. Titles identify the linked version; years marked arXiv are initial posting years, not necessarily conference years. Journal versions are labeled by venue. Foundational readings remain useful even when newer work is available.
Sutton & Barto · 2018↗02A Survey on Large Language Model based Autonomous AgentsFurther · Module 01
Wang et al. · arXiv 2023↗03Proximal Policy Optimization AlgorithmsFurther · Modules 01–02
Schulman et al. · arXiv 2017↗04Training language models to follow instructions with human feedbackCore · Module 02
Ouyang et al. · arXiv 2022↗05Direct Preference Optimization: Your Language Model is Secretly a Reward ModelCore · Module 02
Rafailov et al. · arXiv 2023↗06DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsFurther · Module 02 · GRPO
Shao et al. · arXiv 2024↗07DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learningFurther · Modules 02–03
Guo et al. · Nature 2025↗08Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsFurther · Module 03
Wei et al. · arXiv 2022↗09Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsFurther · Module 03
Yao et al. · arXiv 2023↗10ReAct: Synergizing Reasoning and Acting in Language ModelsCore · Modules 03 & 05
Yao et al. · arXiv 2022↗11Generative Agents: Interactive Simulacra of Human BehaviorCore · Module 04
Park et al. · arXiv 2023↗12Toolformer: Language Models Can Teach Themselves to Use ToolsCore · Module 05
Schick et al. · arXiv 2023↗13AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent ConversationCore · Module 06
Wu et al. · arXiv 2023↗14WebArena: A Realistic Web Environment for Building Autonomous AgentsCore · Module 07
Zhou et al. · arXiv 2023↗15AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsCore · Modules 07 & 11
Debenedetti et al. · arXiv 2024↗16SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Further · Module 08
Jimenez et al. · arXiv 2023↗17SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringCore · Module 08
Yang et al. · arXiv 2024↗18The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoveryCore · Module 09
Lu et al. · arXiv 2024↗19Mastering diverse control tasks through world modelsFurther · Module 10
Hafner et al. · Nature 2025↗20OpenVLA: An Open-Source Vision-Language-Action ModelCore · Module 10
Kim et al. · arXiv 2024↗21Reflexion: Language Agents with Verbal Reinforcement LearningCore · Module 12
Shinn et al. · arXiv 2023↗22AgentBench: Evaluating LLMs as AgentsCore · Module 13
Liu et al. · arXiv 2023↗
Course policies
Expectations and support.
These are planning guidelines. The official Fall 2026 syllabus will supersede this page where they differ. Students are responsible for reviewing UMD’s course-related policies and resources.
Teams & collaboration
Projects use teams of up to three students. Record individual contributions in the midterm and final reports. A self-proposed topic requires written instructor approval; individual quizzes and other designated individual work must be completed independently.
Generative AI use
AI tools may be used only when an assignment permits them. Disclose the model and version, preserve material prompts and tool traces, describe substantive edits, and verify every claim. AI assistance is not permitted on in-class quizzes unless explicitly stated.
Academic integrity
Follow the UMD Code of Academic Integrity and Honor Pledge. Cite external code, models, datasets, papers, prompts, and tools. Do not copy or distribute solutions, fabricate evidence, or present another person’s or system’s work as your own.
Late work
A shared 72-hour late bank applies across homework and written reports. It does not apply to quizzes, live presentations, or other in-person assessments. Once the bank is exhausted, late work is accepted only through an approved accommodation or excused-absence process.
Attendance & absences
Regular attendance and active participation are expected. Report known absences before the schedule-adjustment deadline and unexpected absences as soon as possible through a private ELMS or course-Slack message. Do not post medical documentation publicly.
Responsible experimentation
Use least privilege, sandbox side effects, protect private data, and obtain approval before tests that touch people or external systems. Human-subject research requires appropriate approval before data collection; consequential actions require monitoring and rollback ownership.
Accessibility & support
Students who need accommodations should contact the University’s accessibility service and the instructor early. Please raise barriers to course materials, assessments, or participation as soon as possible so arrangements can be made.
Communication & continuity
ELMS is the official source for announcements, assignments, grades, and urgent changes; course Slack supports discussion and help. If campus operations are interrupted, the continuation plan and any deadline changes will be announced through ELMS.
Planning page last updated September 11, 2026. Material changes after the first class will be dated and announced through ELMS.
CMSC 848N · Fall 2026
Models speak. Agents change state.
Our job is to understand the system in between—and demand evidence for what it can safely do.