ThinMQM (automated translation evaluation, MQM) model and data collection.
Runzhe Zhan
rzzhan
AI & ML interests
None yet
Recent Activity
upvoted a paper about 2 hours ago
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives liked a model 9 days ago
deepseek-ai/DeepSeek-V4-Flash-0731 upvoted a paper 13 days ago
HumanCLAW: Can Vision-Language Models Act Through a Body?