Skip to main content

Yichen Wang

Sep 24, 2026

3:12
Could you explain what VLA models are and why they may be particularly important for the future of flexible gastrointestinal endoscopy?
3:19
Right.
3:20
Of course.
3:20
Uh, so, uh, VLA model is, uh, like you said, actually developed, uh, on top of large language models.
3:27
Um, in large language models, uh, systems takes, uh, text as input and then, uh, predicts the, predicts the next token as output.
3:37
Now, some of the large language models can actually take other modality of information as, as input, such as, uh, videos and pictures, uh, but still output, uh, the next token.
3:46
Uh, we call these models as vision language models.
6:00
What do you see as the biggest technical and clinical barriers that must be overcome before AI-assisted robotic systems can safely participate in advanced therapeutic interventions?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.