Fine-grained robot manipulation demands more than completing a task once. Reliable physical intelligence requires accurate instruction understanding, precise local perception, and stable, constraint-aware action execution. However, most existing benchmarks still summarize performance using binary success rates, making it difficult to determine why a system succeeds, where it fails, and which capability should be improved.
Our workshop brings together researchers in robot learning, manipulation, embodied AI, computer vision, and vision-language-action modeling to rethink how fine-grained robotic capabilities should be evaluated. The workshop is built around MetaFine, a diagnostic meta-evaluation framework that decomposes manipulation competence into three dimensions—understanding, perception, and controlled behavior—and reveals capability-specific limitations hidden by aggregate metrics.
A central component of the workshop is the MetaFine Open Competition. Participating teams will evaluate and improve their systems under a unified diagnostic protocol, with the top teams invited to present their approaches and analyze representative successes and failures during the workshop. Competition results will serve as shared evidence for expert discussions, collaborative failure analysis, breakout problem-solving, and community debate.
Rather than functioning as a sequence of talks, the workshop will actively engage speakers, competition participants, paper authors, and attendees in addressing concrete open questions: How should manipulation failures be attributed? How can robustness and causal bottlenecks be evaluated? How should simulation and real-world evidence be combined? And how can multidimensional capabilities be reported without collapsing them into another misleading single score?
Our goal is to move robot-learning evaluation from ranking to diagnosis, and to establish actionable principles for measuring and improving genuine physical dexterity.