Mimic Robotics
23/7/26Blog

Introducing FLUX-mimic: Scaling Video-Action Models for General Purpose Dexterity

Introducing FLUX-mimic, a next-generation Video-Action Model for general purpose dexterity, developed in partnership with Black Forest Labs.

Introducing FLUX-mimic: Scaling Video-Action Models for General Purpose Dexterity
16/7/26Blog

Solving Dexterity: A Full-Stack Approach

mimic robotics unveils the mimic hand M1 and mimic wearable U1, a full-stack physical AI platform for general-purpose dexterous manipulation.

Solving Dexterity: A Full-Stack Approach
6/1/26Blog

Video-Action Models: Are video model backbones the future of VLAs?

This blog post is about mimic-video, our latest mimic release in which we instantiate a new class of Video-Action Models (VAM), grounding robotic policies in pretrained video models. We argue that video model backbones can be a much more natural choice for robotics foundation model pre-training compared to VLM backbones.

Video-Action Models: Are video model backbones the future of VLAs?
17/12/25arXiv

mimic-video: Video-Action Models

Generalizable robot control beyond VLAs with pretrained Internet-scale video model backbones.

mimic-video: Video-Action Models
17/7/25CoRL

Diffusion for Cross-Embodiment Manipulation

A novel diffusion-based framework to align action spaces across human hands, mimic’s humanoid robotic hands, and other grippers. By training a unified latent action space via contrastive learning, LAD enables a single policy to control diverse robot embodiments, delivering up to a ~25 % lift in manipulation success and unlocking scalable skill transfer across robotic platforms.

Diffusion for Cross-Embodiment Manipulation
1/7/25Blog

Form Follows Data

One of the first questions we are always asked at mimic is: "Why the human hand form factor?" It's a fair question. Indeed, it is not obvious why we would want to focus on this specific morphology, when the design space of grippers allows for so many alternatives, from simpler two finger grippers, to hands more capable than the human, for example with 6 fingers. Here I want to lay out the design philosophy guiding our quest towards solving general-purpose robotic dexterity and what I think the future of the field will look like. As the title already suggests, the reason isn't purely functional: it’s a choice driven by data availability.

Form Follows Data
13/6/25CoRL

mimic-one: a Scalable Model Recipe

A next-generation diffusion-based control architecture built around a newly designed 16‑DoF tendon‑driven humanoid hand. Using glove and VR teleoperation data, MIMIC‑ONE achieves up to 93.3 % task success (even on unseen tasks) by learning smooth, self-correcting motion policies from sensory inputs alone.

mimic-one: a Scalable Model Recipe
8/4/25arXiv

MAPLE: Priors Learned From Videos

In collaboration with ETH Zürich and Microsoft Research, MAPLE introduces a new method to train real-world dexterous robotic hands by learning from large-scale video of human hand interactions. It shows how manipulation priors extracted from egocentric footage enable robots to grasp, manipulate, and generalize across varied tasks.

MAPLE: Priors Learned From Videos
24/10/24Blog

Robotics in the era of the Scaling Hypothesis

This piece is my attempt to collect my thoughts on the current landscape in AI and Robotics, how this informs our approach at mimic in our quest towards solving general-purpose robotic dexterity, and what I think the future of the field will look like.

Robotics in the era of the Scaling Hypothesis