Instructors (TAs)
Fall 2026 · CMSC 848N · University of Maryland
Generative
AI Agents
Foundations and frontiers of systems that reason, plan, act, use tools, and adapt in complex environments.
The course
Build agents whose claims can survive contact with the world.
An agent is more than a model. It is a governed sequential decision system: model, controller, environment, state and memory, tools and protocols, authority policy, evaluator, and people.
We will connect theoretical foundations to modern practice, with an emphasis on algorithmic building blocks, valid evaluation, and open research problems.
CSI 3117 · America/New_York
Probability, optimization, linear algebra, and Python
Diagnosis, implementation, and one personalized research project
Readings, homework, experiments, and team research
Course links and office hours will be shared through ELMS and course Slack
Schedule & slides
Thirteen modules.
Fall classes run August 31–December 11. CMSC 848N meets September 1–December 10; new lecture slide decks appear exactly two days before each 11:00 a.m. lecture. Dates follow the official UMD semester calendar.
No class Tuesday, October 13 (Fall Break). W7-L1 meets Thursday, October 15.
- Module01
Agent-system foundations
Define the complete agent system and reason about sequential experience.
- Module02
Reinforcement learning foundations
Complete the reinforcement-learning foundations needed for later agent training.
Reinforcement Learning for Generative Agents (continued)Continued from prior lecturePre-class reading →Reinforcement Learning for Generative Agents (continued)Continued from prior lecturePre-class reading → - Module03
GRPO and alignment
Connect group-relative reinforcement learning with preference optimization and controlled decoding.
- Module04
Alignment continuation and planning
Complete the alignment material, then study planning and state.
- Module05
Agent harnesses, memory, and tool use
Study how the harness selects context, maintains memory, and executes tool calls.
- Module06
Durable composition
Compose agents into workflows and justified multi-agent teams.
- Module07
Inside a coding-agent harness
Follow a real repair and explain how the harness shapes actions, feedback, and verification.
- Module08
Training coding agents and improving their harnesses
Learn from repair trajectories, then test changes to the runtime around a fixed model.
Sources and reuse notices for the three coding-agent lectures - Module09
Research and scientific agents
Connect sources, claims, hypotheses, experiments, and qualified conclusions.
Deep Research Agents: Search, Evidence, and Reproducible SynthesisPDF release: Sun, Oct 25 · 11 AMPre-class reading →Scientific Discovery Agents: Hypotheses, Experiments, and Valid ClaimsPDF release: Tue, Oct 27 · 11 AMPre-class reading → - Module10
Embodied agents and robotics
Design and evaluate closed-loop world-model and vision-language-action agents.
World Models and Vision-Language-Action AgentsPDF release: Sun, Nov 1 · 11 AMPre-class reading →Data-Efficient Embodied Learning and EvaluationPDF release: Tue, Nov 3 · 11 AMPre-class reading → - Module11
Agent robustness and safety
Threat-model the lifecycle and plan prevention, monitoring, and recovery.
Adversarial Threats Across the Agent LifecyclePDF release: Sun, Nov 8 · 11 AMPre-class reading →Defense in Depth, Monitoring, and Incident ResponsePDF release: Tue, Nov 10 · 11 AMPre-class reading → - Module12
Learning from experience and self-improvement
Govern persistent adaptation with evaluation, lineage, retention, and rollback.
Learning from Experience: Prompts, Skills, Policies, and Continual AdaptationPDF release: Sun, Nov 15 · 11 AMPre-class reading →Self-Modifying and Evolutionary AgentsPDF release: Tue, Nov 17 · 11 AMPre-class reading → - Module13
Evaluation, provenance, and synthesis
Make deployment decisions from reproducible evidence and accountable provenance.
Agent Evaluation and Deployment EvidencePDF release: Sun, Nov 22 · 11 AMPre-class reading →Provenance, Watermarking, and Frontier ChallengesPDF release: Sun, Nov 29 · 11 AMPre-class reading → - Upcoming—
December 3 lecture
The topic and reading assignment will be confirmed.
Topic to be confirmedTopic and readings pendingPre-class reading →
UMD Fall Break
UMD Thanksgiving Recess
Personalized research project presentations
Prepare for each lecture
Pre-class readings.
Start with Read first. Optional readings provide more detail or a different approach. Use the reading focus to guide your preparation.
These links are available ahead of the slide PDFs. Upcoming lecture readings are marked Planned and may be revised as the course develops.
Updated . Titles identify the linked versions; arXiv years are initial posting years. Textbooks, engineering articles, and official specifications are labeled separately.
Agent Systems: The Model Is Not the Agent
Reading focus: Identify the model, memory, planning, and action components of an agent. What information passes between them?
Read first
- A Survey on Large Language Model based Autonomous AgentsPaper · Wang et al. · arXiv 2023
Focus on the agent framework and its main components.
Optional reading (1)
- ReAct: Synergizing Reasoning and Acting in Language ModelsPaper · Yao et al. · arXiv 2022
Reinforcement Learning for Generative Agents
Reading focus: Write down an MDP and distinguish a policy, a value function, and a Bellman equation.
Read first
- Reinforcement Learning: An Introduction (2nd edition)Textbook · Sutton & Barto · 2018
Chapters 3–4: finite Markov decision processes and dynamic programming.
Reinforcement Learning for Generative Agents (continued)
Reading focus: For the RL continuation, compare a sampled return with a bootstrapped value target.
Read first
- Reinforcement Learning: An Introduction (2nd edition)Textbook · Sutton & Barto · 2018
Chapter 6: temporal-difference learning, especially TD prediction, SARSA, and Q-learning.
Optional reading (1)
- Reinforcement Learning: An Introduction (2nd edition)Textbook · Sutton & Barto · 2018
Chapter 9: on-policy prediction with function approximation.
Reinforcement Learning for Generative Agents (continued)
Reading focus: For the RL continuation, follow the policy-gradient derivation and explain the role of a baseline.
Read first
- Reinforcement Learning: An Introduction (2nd edition)Textbook · Sutton & Barto · 2018
Chapter 13: policy-gradient methods and actor–critic methods.
Optional reading (1)
- Proximal Policy Optimization AlgorithmsPaper · Schulman et al. · arXiv 2017
Focus on the clipped surrogate objective; we will return to it when studying GRPO.
DeepSeek-R1, GRPO, and Agentic Reinforcement Learning
Reading focus: Compare PPO and GRPO: how are advantages estimated, and what does the reference policy do?
Read first
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsPaper · Shao et al. · arXiv 2024
Focus on the GRPO objective and group-relative advantage estimates.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningPaper · DeepSeek-AI et al. · arXiv 2025
Focus on the training stages, reward design, and evaluation.
Optional reading (1)
- Proximal Policy Optimization AlgorithmsPaper · Schulman et al. · arXiv 2017
Alignment from DPO to Controlled Decoding
Reading focus: Derive the DPO loss from KL-regularized reward maximization, then compare parameter updates with control during decoding.
Read first
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelPaper · Rafailov et al. · arXiv 2023
Follow the reward–policy relationship and preference-likelihood derivation.
- Controlled Decoding from Language ModelsPaper · Mudgal et al. · ICML 2024
Focus on the tokenwise objective and prefix value function.
Optional reading (4)
- Training language models to follow instructions with human feedbackPaper · Ouyang et al. · arXiv 2022
- Transfer Q Star: Principled Decoding for LLM AlignmentPaper · Chakraborty et al. · arXiv 2024
- GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time AlignmentPaper · Xu et al. · arXiv 2024
- Collab: Controlled Decoding using Mixture of Agents for LLM AlignmentPaper · Chakraborty et al. · arXiv 2025
Alignment from DPO to Controlled Decoding (continued)
Reading focus: Continue the September 17 alignment material using the same slides and readings.
Read first
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelPaper · Rafailov et al. · arXiv 2023
Revisit the reward-policy relationship and preference-likelihood derivation.
- Controlled Decoding from Language ModelsPaper · Mudgal et al. · ICML 2024
Revisit the tokenwise objective and prefix value function.
Optional reading (4)
- Training language models to follow instructions with human feedbackPaper · Ouyang et al. · arXiv 2022
- Transfer Q Star: Principled Decoding for LLM AlignmentPaper · Chakraborty et al. · arXiv 2024
- GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time AlignmentPaper · Xu et al. · arXiv 2024
- Collab: Controlled Decoding using Mixture of Agents for LLM AlignmentPaper · Chakraborty et al. · arXiv 2025
Planning and State
Reading focus: What state does a planner need, how does it predict an action's outcome, and when should it replan?
Read first
- ReAct: Synergizing Reasoning and Acting in Language ModelsPaper · Yao et al. · arXiv 2022
- Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web AgentsPaper · Gu et al. · arXiv 2024
Focus on how the agent simulates and compares candidate actions.
Optional reading (3)
- LLM+P: Empowering Large Language Models with Optimal Planning ProficiencyPaper · Liu et al. · arXiv 2023
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsPaper · Xie et al. · arXiv 2024
- Tree Search for Language Model AgentsPaper · Koh et al. · arXiv 2024
Memory and Context Engineering
Reading focus: Trace how the agent harness builds context for each model call. Then compare storing, retrieving, and summarizing experience. What evidence can each step lose?
Read first
- Generative Agents: Interactive Simulacra of Human BehaviorPaper · Park et al. · arXiv 2023
Focus on the memory stream, retrieval, reflection, and planning.
Optional reading (12)
- Effective context engineering for AI agentsEngineering article · Anthropic · 2025
Focus on compaction and structured notes; this is a practitioner article.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPaper · Lewis et al. · arXiv 2020
- Voyager: An Open-Ended Embodied Agent with Large Language ModelsPaper · Wang et al. · arXiv 2023
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsPaper · Gutiérrez et al. · arXiv 2024
- Reflexion: Language Agents with Verbal Reinforcement LearningPaper · Shinn et al. · arXiv 2023
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryPaper · Wu et al. · arXiv 2024
- Cognitive Architectures for Language AgentsPaper · Sumers et al. · arXiv 2023
Cited in the slides: CoALA.
- Referral Augmentation for Zero-Shot Information RetrievalPaper · Tang et al. · arXiv 2023
Cited in the slides.
- FireAct: Toward Language Agent Fine-tuningPaper · Chen et al. · arXiv 2023
Cited in the slides.
- Large Language Models Are Human-Level Prompt EngineersPaper · Zhou et al. · arXiv 2022
Cited in the slides: Automatic Prompt Engineer (APE).
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringPaper · Yang et al. · arXiv 2024
Cited in the slides.
- Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better TogetherPaper · Soylu et al. · arXiv 2024
Cited in the slides: BetterTogether.
Tool Use and Action Interfaces
Reading focus: How does an agent learn when to call a tool, and what must the tool interface make explicit?
Read first
- Toolformer: Language Models Can Teach Themselves to Use ToolsPaper · Schick et al. · arXiv 2023
Optional reading (6)
- Model Context Protocol specificationOfficial specification · MCP contributors · version 2025-11-25
Read the overview, tools, and security sections of the version used in the lecture draft.
- ReAct: Synergizing Reasoning and Acting in Language ModelsPaper · Yao et al. · arXiv 2022
- The Protection of Information in Computer SystemsPaper · Saltzer & Schroeder · 1975
Optional background for tool permissions: Section I, least privilege and complete mediation.
- Executable Code Actions Elicit Better LLM AgentsPaper · Wang et al. · ICML 2024
Cited in the slides: CodeAct.
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language ModelsPaper · Dong et al. · arXiv 2024
Cited in the slides.
- Defeating Prompt Injections by DesignPaper · Debenedetti et al. · arXiv 2025
Cited in the slides: CaMeL.
Agentic Workflows and Orchestration
Reading focus: Describe a workflow's branches, shared state, and recovery after a partially completed operation.
Read first
- Workflow PatternsPaper · van der Aalst et al. · Distributed and Parallel Databases, 2003
Focus on sequence, parallel split, synchronization, choice, and merge patterns. Publisher access may require a university login.
Optional reading (3)
- Building effective agentsEngineering article · Anthropic · 2024
Compare the workflow examples and their tradeoffs.
- SagasPaper · Garcia-Molina & Salem · 1987 technical report
- AFlow: Automating Agentic Workflow GenerationPaper · Zhang et al. · arXiv 2024
Multi-Agent Systems
Reading focus: When does dividing a task among agents help, and how can coordination fail?
Read first
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent ConversationPaper · Wu et al. · arXiv 2023
Optional reading (1)
- Why Do Multi-Agent LLM Systems Fail?Paper · Cemri et al. · arXiv 2025
Inside a Coding-Agent Harness
Reading focus: Follow a real repair from a failed reproduction through rejected edits to a submitted patch. How do the harness's tools, feedback, and stopping rules affect what the model can do and what the result verifies?
Read first
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringPaper · Yang et al. · arXiv 2024 · v3
Focus on the search and editing interfaces, the editor ablation, and the lint-gate limitation. Separate a reported benchmark result from the evidence in one repair trace.
- mini-SWE-agent: Agent control flowDocumentation · mini-SWE-agent project · official documentation
Follow the control-flow diagram and the run, step, query, and execute_actions methods.
Optional reading (3)
- mini-SWE-agent: Default agent implementationCode · SWE-agent project · commit 04d809ce
Inspect the pinned implementation of the execution loop and its error and stopping paths.
- SWE-agent: Recorded pydicom repair trajectoryExecution record · SWE-agent project · commit 3ea751c0
Follow the pydicom repair used in class. Identify what the local reproduction and final submission establish.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Paper · Jimenez et al. · arXiv 2023
Background on task construction and the independent evaluation of a returned patch.
Training Coding Agents
Reading focus: Turn executable repair tasks and interaction records into training data. How do trajectory selection, token loss masks, and verifier rewards determine what the model learns?
Read first
- SWE-smith: Scaling Data for Software Engineering AgentsPaper · Yang et al. · arXiv 2025 · v2
Follow task generation, validation, trajectory collection, and the reported training setup. Ask what each test and retained trajectory establishes.
- SWE-smith: Train SWE-agentsDocumentation · SWE-smith project · official documentation
Trace the steps from collected trajectories to supervised fine-tuning. Identify where successful-run filtering and training configuration enter.
Optional reading (4)
- SWE-smith: Trajectory collection and conversionCode · SWE-smith project · commit 9b74ac0
Distinguish attaching a resolved label from filtering out unsuccessful records.
- SWE-smith: Qwen 32B training configurationCode · SWE-smith project · commit 9b74ac0
Inspect the dataset and train_on_input settings. Determine which token labels would be needed to verify the loss mask.
- SWE-Universe: Scale Real-World Verifiable Environments to MillionsPaper · Chen et al. · arXiv 2026 · v1
Compare how tasks and verifiers are constructed and checked. The lecture's GRPO derivation is a separate course formulation.
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsPaper · Shao et al. · arXiv 2024 · v3
Review Section 4.1 for the group-relative advantage and clipped objective before the multi-turn coding example.
Optimizing Coding-Agent Harnesses
Reading focus: Hold the model fixed and propose a change to its runtime. What evidence shows that the change helps on new tasks after accounting for repeated selection, feedback, and compute?
Read first
- Self-Harness: Harnesses That Improve ThemselvesPaper · Zhang et al. · arXiv 2026 · v3
Read the search procedure and Algorithm 1's acceptance rule. Track which evaluation results are reused to retain an edit.
- Rethinking the Evaluation of Harness Evolution for AgentsPaper · Wang et al. · arXiv 2026 · v4
Focus on matched-feedback baselines and evaluation on unseen tasks. Identify which controls support a generalization claim.
Optional reading (4)
- RRSI: Regularized Recursive Self-Improvement of Agent HarnessesPaper · Xia et al. · arXiv 2026 · v3
Examine how proposal and selection restrictions aim to control overfitting and resource growth.
- RRSI: Candidate selection implementationCode · Google Research · commit be50316
Inspect the acceptance conditions, including the score floor, cost rule, and domain guards.
- An Empirical Study of Harness Design for Coding AgentsPaper · Fan et al. · arXiv 2026 · v1
Compare the effects of planning, action interfaces, and context management across models and budgets.
- AutoCompact: Learning When to Compact Context in Long-Horizon Coding AgentsPaper · Zhang et al. · arXiv 2026 · v1
Explain why learning when and how to compact context changes the policy, even though compaction operates at runtime.
Deep Research Agents: Search, Evidence, and Reproducible Synthesis
Reading focus: How should a research agent gather sources, support its claims, and check citation quality?
Read first
- Assisting in Writing Wikipedia-like Articles From Scratch with Large Language ModelsPaper · Shao et al. · arXiv 2024
The STORM paper: focus on question generation, source gathering, and article organization.
Optional reading (2)
- Enabling Large Language Models to Generate Text with CitationsPaper · Gao et al. · arXiv 2023
- GAIA: a benchmark for General AI AssistantsPaper · Mialon et al. · arXiv 2023
Scientific Discovery Agents: Hypotheses, Experiments, and Valid Claims
Reading focus: Distinguish proposing a hypothesis, running an experiment, and establishing a scientific result.
Read first
- The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoveryPaper · Lu et al. · arXiv 2024
Optional reading (2)
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryPaper · Chen et al. · arXiv 2024
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringPaper · Chan et al. · arXiv 2024
World Models and Vision-Language-Action Agents
Reading focus: Compare learning a world model with learning a policy conditioned on vision and language.
Read first
- Mastering Diverse Domains through World ModelsPaper · Hafner et al. · arXiv 2023
The DreamerV3 paper: focus on the learned world model and policy learning.
- OpenVLA: An Open-Source Vision-Language-Action ModelPaper · Kim et al. · arXiv 2024
Data-Efficient Embodied Learning and Evaluation
Reading focus: What transfers across robots and tasks, and what evidence supports transfer to a new setting?
Read first
- Open X-Embodiment: Robotic Learning Datasets and RT-X ModelsPaper · Open X-Embodiment Collaboration · arXiv 2023
Optional reading (2)
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot LearningPaper · Liu et al. · arXiv 2023
- Evaluating Real-World Robot Manipulation Policies in SimulationPaper · Li et al. · arXiv 2024
The SIMPLER paper: examine the relationship between simulation and real-world evaluation.
Adversarial Threats Across the Agent Lifecycle
Reading focus: How can poisoned memory or retrieved knowledge change an agent's behavior? State the attacker's access and objective.
Read first
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesPaper · Chen et al. · arXiv 2024
Optional reading (1)
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language ModelsPaper · Zou et al. · arXiv 2024
Defense in Depth, Monitoring, and Incident Response
Reading focus: What can a monitor observe, which actions can it stop, and what assumptions does the safety argument require?
Read first
- AI Control: Improving Safety Despite Intentional SubversionPaper · Greenblatt et al. · arXiv 2023
Optional reading (2)
- Defeating Prompt Injections by DesignPaper · Debenedetti et al. · arXiv 2025
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsPaper · Debenedetti et al. · arXiv 2024
Learning from Experience: Prompts, Skills, Policies, and Continual Adaptation
Reading focus: What persists after feedback: a written reflection, a reusable skill, or a change to the model?
Read first
- Reflexion: Language Agents with Verbal Reinforcement LearningPaper · Shinn et al. · arXiv 2023
Optional reading (2)
- Self-Refine: Iterative Refinement with Self-FeedbackPaper · Madaan et al. · arXiv 2023
- Voyager: An Open-Ended Embodied Agent with Large Language ModelsPaper · Wang et al. · arXiv 2023
Self-Modifying and Evolutionary Agents
Reading focus: How are proposed agent modifications generated, evaluated, and selected? How would you detect overfitting to the development tasks?
Read first
- Automated Design of Agentic SystemsPaper · Hu et al. · arXiv 2024
Optional reading (2)
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsPaper · Zhang et al. · arXiv 2025
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code GenerationPaper · Zelikman et al. · arXiv 2023
Agent Evaluation and Deployment Evidence
Reading focus: What do repeated trials, changed conditions, and compute budgets reveal that one average success rate misses?
Read first
- Towards a Science of AI Agent ReliabilityPaper · Rabanser et al. · arXiv 2026
Read as a 2026 preprint; focus on the proposed reliability dimensions and their measurement.
Optional reading (1)
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human expertsPaper · Wijk et al. · arXiv 2024
Provenance, Watermarking, and Frontier Challenges
Reading focus: State the assumptions behind watermark detection and identify what a positive or negative test can establish.
Read first
- A Watermark for Large Language ModelsPaper · Kirchenbauer et al. · arXiv 2023
Optional reading (1)
- Can AI-Generated Text be Reliably Detected?Paper · Sadasivan et al. · arXiv 2023
Topic to be confirmed
The topic and reading assignment will be confirmed. No pre-class reading is assigned yet.
Personalized research project
Choose one project. Make it your own.
Professor Huang is designing a new catalogue of ambitious research candidates for the class. Each begins with a real open problem and a credible path toward a publishable contribution. Detailed briefs are available to enrolled students with the course password.
Protected project catalogue
Unlock the research projects.
The detailed catalogue is encrypted. Enter the password shared by Professor Huang to review the candidate questions, technical hypotheses, evidence requirements, and starting papers.
How project selection works
Review
Read the protected catalogue and identify several projects that fit your preparation and interests.
Rank
Rank three choices and propose one way to personalize your preferred direction. The planned deadline is Friday, September 11 at 5:00 p.m.
Form a team
Projects are completed in teams of up to three students. You may indicate preferred teammates when submitting your rankings.
Confirm
Final assignments will balance preparation, project demand, and available compute. Submission instructions will be posted in ELMS.
Fall 2026 candidate catalogue
Open a brief. Find the question you want to sharpen.
Rank several candidates and propose one personalization. Final scopes will reflect team preparation, available compute, and overlap across the class.
What each research project description includes
Motivation & gap
Why the problem matters and what unresolved gap makes it worth pursuing.
Question & novelty
A concrete research question and the seed of a new, defensible contribution.
Closest prior work
Starting references, relevant systems, and the work a new result must improve upon.
Research hypothesis
A technically credible first route, with space for the team to develop its own approach.
Data & baselines
Datasets, environments, comparison points, and practical starting assets.
Conference-level evidence
Metrics, budgets, strong baselines, ablations, and failure analyses for a valid claim.
Research artifacts
A conference-style paper, reproducible code, evidence, and a concise presentation.
Milestones
- Weeks 1–2Rank projects and form teams
Submit project preferences, propose one personalization, and establish a shared workspace with a team of up to three students.
- Week 4Proposal
Submit the task contract, benchmark or environment, and one-page plan.
- Week 6Baseline
Deliver a runnable baseline and the first reproducible evidence bundle.
- Week 8Midterm report
Report preliminary results and a concrete failure taxonomy.
- Week 10Freeze evaluation
Lock held-out tasks, metrics, budgets, and contamination controls.
- Week 12Artifact audit
Peer review, regression check, security test, and release rehearsal.
- Dec 8 & 10Present and submit
Deliver the final presentation, paper-style report, runnable code, and supporting evidence.
Assessment
How your work will be assessed.
The assessment structure follows the Fall 2025 syllabus. Dates below are planned; the final Fall 2026 syllabus will confirm them, grade thresholds, and whether plus/minus grading is used.
In-class quizzes
Impromptu short quizzes during lectures.
Sep 1–Dec 3Homework
Three assignments covering the course material.
Sep 30 · Oct 30 · Nov 30Midterm report
Project progress, a literature review, and baseline results.
Oct 15Final presentation
A concise presentation of the project, results, and limitations.
Dec 8 & 10Final report and artifacts
A research or engineering report, runnable code, and supporting evidence.
Dec 11Project assessment
Project grading rubrics.
Proposed Fall 2026 rubrics, subject to confirmation in the final syllabus. Research and engineering tracks share expectations for technical depth, sound evaluation, and reproducibility.
Teams select a track and agree on scope with the instructor at the proposal milestone. Hybrid projects use one agreed rubric. The midterm package corresponds to the midterm report; final artifacts include the final report and supporting materials.
Each table totals 100 points within its component. The two final-artifact rubrics are alternatives. In-class quizzes (10%) and homework (20%) remain separate assessments.
Midterm package
30% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Problem and scope | 20 | A precise question or use case, motivated by a concrete limitation; clearly defined success criteria and a feasible semester scope. |
| Related work and baseline | 25 | Accurate discussion of the closest approaches, verified citations, and a runnable baseline appropriate to the question. |
| Preliminary evidence | 30 | An end-to-end experiment or prototype, interpretable measurements, and an analysis of failures and uncertainty. |
| Evaluation plan and milestones | 25 | Held-out tasks, meaningful metrics, fair comparisons, resource budgets, major risks, and a credible plan for the remaining work. |
Final presentation
20% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Problem and approach | 20 | Clear motivation, necessary background, and an understandable explanation of the technical choices. |
| Evidence and interpretation | 40 | Results against appropriate baselines, readable figures, and conclusions supported by measurements. A demo supplements the evidence. |
| Limitations and discussion | 20 | Thoughtful treatment of failure cases, alternative explanations, uncertainty, and questions from the audience. |
| Communication and contributions | 20 | A coherent presentation within the allotted time, legible slides, and clearly identified contributions from each team member. |
Final artifacts: research track
20% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Research question and contribution | 20 | A precise hypothesis and a clear account of what the work adds relative to the closest prior work. |
| Experimental rigor | 35 | Strong baselines, controlled comparisons, relevant ablations, held-out evaluation, and uncertainty estimates where appropriate. |
| Technical correctness and analysis | 25 | Correct methods and mathematics, defensible conclusions, and substantive analysis of failures and competing explanations. |
| Reproducibility and reporting | 20 | A clear paper-style report, runnable code, documented configurations and dependencies, and sufficient details to reproduce the main results. |
Final artifacts: engineering track
20% of course grade
| Criterion | Points | What earns full credit |
|---|---|---|
| Correctness and reliability | 35 | An end-to-end system meeting agreed requirements, with automated tests, failure handling, and appropriate security and privacy protections. |
| Evaluation | 25 | Comparison with a simple baseline under documented workloads, measuring relevant quality, latency, cost, or usability outcomes. |
| System design | 20 | Well-justified interfaces and architectural choices, explicit tradeoffs, and a maintainable implementation. |
| Reproducibility and documentation | 20 | Reliable setup instructions, a repeatable demo, test commands, configurations, and clearly documented limitations. |
What to submit
- Midterm: a short report, runnable code, preliminary results, and an updated experimental plan.
- Final presentation: slides explaining the problem, approach, results, limitations, and individual contributions.
- Final artifacts: the research or engineering report, code, setup and evaluation commands, configurations, and supporting evidence. Engineering submissions also include automated tests and a repeatable demo.
Scoring and fairness
Score each criterion on a 0–4 scale, allowing half-points: 4 fully meets the criterion with convincing evidence; 3 largely meets it with minor gaps; 2 partially meets it with substantial gaps; 1 shows limited progress; 0 is missing or not assessable.
Criterion points = allocated points × score ÷ 4. For example, 80/100 on final artifacts contributes 16 of the 20 available course points.
- Negative results can earn full credit when the investigation is rigorous and informative.
- Conference acceptance, a novel algorithm, expensive compute, and positive results are not grading requirements.
- Compare methods under fair resource budgets and explain unavoidable differences.
- Disclose reused code, datasets, assets, AI assistance, and each member’s contributions.
- Use synthetic or approved data for privacy-sensitive projects. Physical experiments require appropriate safety controls.
- Agree on material scope changes before final evaluation.
- Shared artifacts receive a team grade. Individual adjustments require documented contributions and demonstrated understanding, with an opportunity to clarify discrepancies.
Mathematical prerequisites
Review the foundations.
Use these reviews throughout the semester. For the papers assigned to individual lectures, see the pre-class reading list.
Week 1 · September 1 & 3
Refresh the mathematics behind this week’s lectures.
Review the topics you need before working through the reinforcement-learning definitions, Bellman equations, and policy-improvement derivations.
PDF↗02Linear Algebra ReviewPrerequisite · Calculus & linear algebra
PDF↗03Advanced Linear Algebra ReviewPrerequisite · Calculus & linear algebra
PDF↗04Convex Analysis ReviewPrerequisite · Optimization
PDF↗05Optimization ReviewPrerequisite · Optimization
PDF↗06Probability ReviewPrerequisite · Probability & statistics
PDF↗
Course policies
Expectations and support.
These are planning guidelines. The official Fall 2026 syllabus will supersede this page where they differ. Students are responsible for reviewing UMD’s course-related policies and resources.
Teams & collaboration
Projects use teams of up to three students. Record individual contributions in the midterm and final reports. A self-proposed topic requires written instructor approval; individual quizzes and other designated individual work must be completed independently.
Generative AI use
AI tools may be used only when an assignment permits them. Disclose the model and version, preserve material prompts and tool traces, describe substantive edits, and verify every claim. AI assistance is not permitted on in-class quizzes unless explicitly stated.
Academic integrity
Follow the UMD Code of Academic Integrity and Honor Pledge. Cite external code, models, datasets, papers, prompts, and tools. Do not copy or distribute solutions, fabricate evidence, or present another person’s or system’s work as your own.
Late work
A shared 72-hour late bank applies across homework and written reports. It does not apply to quizzes, live presentations, or other in-person assessments. Once the bank is exhausted, late work is accepted only through an approved accommodation or excused-absence process.
Attendance & absences
Regular attendance and active participation are expected. Report known absences before the schedule-adjustment deadline and unexpected absences as soon as possible through a private ELMS or course-Slack message. Do not post medical documentation publicly.
Responsible experimentation
Use least privilege, sandbox side effects, protect private data, and obtain approval before tests that touch people or external systems. Human-subject research requires appropriate approval before data collection; consequential actions require monitoring and rollback ownership.
Accessibility & support
Students who need accommodations should contact the University’s accessibility service and the instructor early. Please raise barriers to course materials, assessments, or participation as soon as possible so arrangements can be made.
Communication & continuity
ELMS is the official source for announcements, assignments, grades, and urgent changes; course Slack supports discussion and help. If campus operations are interrupted, the continuation plan and any deadline changes will be announced through ELMS.
Planning page last updated October 6, 2026. Material changes after the first class will be dated and announced through ELMS.
CMSC 848N · Fall 2026
Models speak. Agents change state.
Our job is to understand the system in between—and demand evidence for what it can safely do.