WeTalkRobots · About

Reconstructing Hands in 3D with Transformers

Learning from Humans · 08/12/2023

Technical Analysis presents an approach that can reconstruct hands in 3D from monocular input. High-level Analogy: Imagine you're trying to sculpt a lifelike 3D hand from just a single photograph. The old way (State of the Art) was like having a reasonably skilled sculptor who had only studied hands in perfect, well-lit studio conditions and used a limited set of tools. They could do a decent job with clear, simple photos, but if the hand was blurry, partly hidden by an object, or in a complicated, natural pose (like holding a wrench or stirring a pot), their sculpture would often look awkward or incorrect. Their understanding was too narrow. The new way with HaMeR is like giving that sculptor two major upgrades: An enormous library of hand photos (Scaled-up Data): Instead of just studio shots, they now have millions of photos of hands from every angle, doing every imaginable activity in real life – cooking, typing, playing sports, some blurry, some clear, some holding objects, some even interacting with other hands. They've also painstakingly created a new, specialized album (HInt) of particularly challenging, 'in-the-wild' hand photos, even marking which parts of the hand were hidden or difficult to see. A super-advanced, highly adaptable assistant (Large Vision Transformer Model): This assistant isn't just following simple rules; it's a powerful AI that can learn incredibly intricate patterns from all those millions of diverse photos. It observes how real hands actually…

Continue to interactive post