Sam Family

SAM stands for segmentation everything. Up to now, there are three models in this big family: SAM1, SAM2 and SAM3. I’ll give a short review about each of them. Also, I’ll introduce some variants and extensions of SAM, for example, applying stronger memory module for better tracking. Finally, I simply go through the metrics used in VOT (video object tracking) and VOS (video object segementation), then simply talk about how we apply them into our task and data. ...

July 30, 2026 | 1793 words | Author: Tan Ke

A Review of Robbyant’s Early-2026 Work

Robbyant is a company under Ant Group, dedicated to building the foundational platform for Embodied AI, bridging the gap between digital intelligence and the physical world. Since the company is still relatively new, I want to quickly review its recent work. In particular, I will study four embodied intelligence model models: spatial perception model, VLA model, world model, and video action model. This diagram in the homepage of Robbyant reflects the vision for embodied intelligence: starting from sensory input, the system first builds spatial intelligence to understand the physical world, then relies on an action model to make decisions and interact with the environment, and finally improves through environmental reward. ...

March 16, 2026 | 3594 words | Author: Tan Ke

CLIP

Paper Reading Notes: “Learning Transferable Visual Models From Natural Language Supervision”
January 1, 2026 | 887 words | Author: Tan Ke

OpenVLA

Paper Reading Notes: “OpenVLA: An Open-Source Vision-Language-Action Model”
December 12, 2025 | 312 words | Author: Tan Ke