NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 3 days ago • 20
NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms Paper • 2606.20408 • Published Jul 6 • 2
NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms Paper • 2606.20408 • Published Jul 6 • 2
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 3 days ago • 20
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 3 days ago • 20
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models Paper • 2508.03483 • Published Jun 17 • 1
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards Paper • 2607.00218 • Published Jun 30 • 1
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards Paper • 2607.00218 • Published Jun 30 • 1
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models Paper • 2508.03483 • Published Jun 17 • 1