Research infrastructure

Data, benchmarks & evaluation.

Furong Lab builds the evidence layer for intelligent systems: datasets that expose meaningful structure, benchmarks that test consequential capabilities, and evaluation environments that reveal where models fail.

20Open resources
16Datasets
20Benchmarks
3Scientific questions

Complete collection

Find the right data or stress test.

Search by capability, modality, or resource name, then follow the direct links to datasets, leaderboards, code, and papers.

20 resources
Representative figure for MomaGraph-Scenes and MomaGraph-Bench2026
Embodied systemsWorld models

MomaGraph-Scenes and MomaGraph-Bench

DatasetBenchmarkVisionRobotics

Task-driven spatial-functional scene graphs and a six-capability evaluation suite for embodied planning and scene understanding.

Scale 1,050 task-oriented subgraphs; 6,278 multiview images; 350+ household scenes

Representative figure for Sequential Embodied Question Answering2026
Embodied systemsWorld models

Sequential Embodied Question Answering

BenchmarkVisionLanguage

An evaluation setting for persistent visual-semantic memory across multiple questions and continuous embodied operation.

Representative figure for SoundnessBench2026
ReasoningReasoning as control

SoundnessBench

DatasetBenchmarkTextScientific reasoning

Tests whether AI research agents can distinguish methodologically sound proposals from plausible but flawed research ideas.

Scale 1,099 machine-learning research proposals

Representative figure for TraceForge-123K and TraceGen Benchmark2026
Embodied systemsWorld models

TraceForge-123K and TraceGen Benchmark

DatasetBenchmarkVideo3D traces

A cross-embodiment corpus and held-out evaluation suite for learning transferable world models from human and robot videos.

Scale 123K videos; 1.8M observation-trace-language triplets; five evaluation environments

Representative figure for TSRBench2026
ReasoningReasoning as control

TSRBench

DatasetBenchmarkTime seriesVision

A multi-task, multimodal time-series reasoning benchmark for evaluating generalist models across heterogeneous temporal tasks.

Representative figure for μ₀ Evaluation Dataset2026
Embodied systemsWorld models

μ₀ Evaluation Dataset

DatasetBenchmarkVision3D traces

Released evaluation episodes, normalization statistics, and model artifacts for transferable 3D interaction-trace prediction.

Scale Evaluation episodes spanning human and robot embodiments

Representative figure for Fictitious Facial Identity Dataset2025
Trust and safetyTrustworthy self-improvement

Fictitious Facial Identity Dataset

DatasetBenchmarkFacesVision-language

A controlled dataset for evaluating whether vision-language models can reliably unlearn fictitious facial identities.

Representative figure for MORSE-5002025
Embodied systemsWorld models

MORSE-500

DatasetBenchmarkVideoLanguage

A programmatically controllable video benchmark for abstract, physical, planning, spatial, and temporal multimodal reasoning.

Scale 500 scripted videos with controllable physical and temporal structure

Representative figure for PropensityBench2025
Trust and safetyTrustworthy self-improvement

PropensityBench

BenchmarkPlatformTextAI agents

Agentic red-teaming environments surface latent behavioral risks that may remain hidden in single-turn safety tests.

Representative figure for ROVER2025
ReasoningWorld models

ROVER

DatasetBenchmarkImageVideo

Benchmarks reciprocal reasoning between multimodal understanding and generation across image, video, audio, and 3D.

Representative figure for TrustGen2025
Trust and safetyTrustworthy self-improvement

TrustGen

BenchmarkPlatformTextImages

A dynamic benchmarking platform for evaluating trustworthiness across generative language, image, and vision-language models.

Representative figure for Zebra-CoT2025
ReasoningReasoning as control

Zebra-CoT

DatasetBenchmarkVisionLanguage

A dataset for teaching and evaluating interleaved vision-language chains of thought instead of text-only explanations.

Representative figure for AutoHallusion2024
Trust and safetyTrustworthy self-improvement

AutoHallusion

DatasetBenchmarkPlatformVisionLanguage

Automatically generates diverse visual-reasoning stress tests for diagnosing hallucination failures in vision-language models.

Representative figure for Easy2Hard-Bench2024
ReasoningReasoning as control

Easy2Hard-Bench

DatasetBenchmarkTextCode

Standardized continuous difficulty labels for profiling language-model performance and easy-to-hard generalization.

Representative figure for Erasing the Invisible2024
Trust and safetyTrustworthy self-improvement

Erasing the Invisible

DatasetBenchmarkCompetitionImagesWatermarking

A NeurIPS competition, dataset, and evaluation toolkit for stress-testing image watermarks under black-box and beige-box attacks.

Representative figure for HallusionBench2024
Trust and safetyTrustworthy self-improvement

HallusionBench

DatasetBenchmarkVisionLanguage

A diagnostic benchmark that disentangles language hallucination from visual illusion in large vision-language models.

Representative figure for Mementos2024
ReasoningWorld models

Mementos

DatasetBenchmarkImage sequencesLanguage

Evaluates whether multimodal language models can reason over coherent image sequences rather than isolated frames.

Representative figure for PHTest2024
Trust and safetyTrustworthy self-improvement

PHTest

DatasetBenchmarkTextSafety

Pseudo-harmful prompts for measuring false refusals and the tradeoff between safety alignment and model helpfulness.

Representative figure for WAVES2024
Trust and safetyTrustworthy self-improvement

WAVES

BenchmarkPlatformImagesWatermarking

A standardized benchmark and stress-testing toolkit for image-watermark detection and identification under diverse attacks.

Representative figure for Easy-to-Hard Generalization Datasets2021
ReasoningReasoning as control

Easy-to-Hard Generalization Datasets

DatasetBenchmarkImagesClassification

Controlled image-classification datasets for studying how models generalize from easy training examples to harder test examples.

No matching resources

Try a broader keyword or clear one of the filters.

Evaluation philosophy

Measure behavior, not just average accuracy.

01

Expose structure

Datasets make states, difficulty, interactions, and failure conditions explicit enough to study.

02

Test transfer

Evaluation asks whether systems generalize across environments, embodiments, modalities, and levels of difficulty.

03

Stress the boundary

Agentic and adversarial tests probe behavior beyond comfortable average-case distributions.