You will stare down death one day. You will close your eyes, and leave your body, dissipating into the blackness. You will breathe your last breath, and be no more — at least, here.

What awaits you? I don’t know. Humble as I can say, despite all my knowledge in the occult…

I truly don’t know.

And yet, this is what makes it so beautiful.

What bravery, what gall, what courage you must muster, my fellow brother. To stare this undefeated, omnipotent, unavoidable conclusion right in the eye. To unwaver yourself before it’s grip.
What better poem could ever be written? What better song? What better act of faith to the Lord of all, that you should stand there in awe, in fear, and yet, see it through

I think about this too much, but it cannot be any other way. God hath stilled the finiteness of life within
The unspeakable drive, this illogical urge — to be.

Ancestors were right. Great men were right. Life is a stage, and what you do here echoes long after into infinities halls. Mortal men dream of being gods, when in fact, they already are. Only a degree of separation lies between
A belief, in themselves.

Dance your dance. Move to the rhythm you feel within. Chase beauty like a star racing across the sky, and when your time comes, face it head on, and God will carry you home.

That I can promise.

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like