1/ We just concluded our #ZohoDay2021 online event for industry analysts, with record participation.🙏

We hosted this on Zoho's own event platform with our home grown audio-video framework, a combination of @ZohoBackstage, @ZohoShowTime and @zohomeeting.

I will explain this.

2/ One year ago, our audio-video frameworks were not ready. As the pandemic struck, we realized we were behind in AV tech.

In March 2020, we redoubled our R&D efforts. I told our engineers we will invest in this however long it takes and produce a world-class platform.
3/ We transferred engineers from other projects to the AV teams. At first the progress was frustratingly slow, as it often happens in R&D, with frequent setbacks. I told our engineers to persist.

We used our own products and diligently reported issues encountered to engineers.
4/ Then by Oct 2020, we started to improve and we could see visible progress. Yet, running an online event entirely on our platform was a dream.

Today, we have concluded a successful event, due to painstaking efforts from our engineers and our events team who stood by them.🙏
5/ This is a classic example of R&D that I often mention. We need a lot more of this across the industrial landscape in India. The private sector must get into R&D with passion and religious devotion.

Company leadership must patiently stand by engineers so they can excel.
6/ We still have work to do in our AV frameworks but we have come a long way in the past year and I am sure our engineers, none of whom had a background in AV before Zoho, will take our products to world class soon. My job is is to ensure I strongly support and nurture them.🙏

More from Sridhar Vembu

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like