SREE KALLIL BHAGAWATHI MANDIR, METHALA (Ker)

Kallil in Malayalam means 'in stone'.
Bhagavathi’s prathishta is in a cave beneath a massive rock that stands without a support on the ground. It appears as if the huge boulder is mysteriously suspended in the air.
@iamkamyabuch

Once, people gathering forest resources in the dense forest saw a beautiful lady toying with huge stones. The lady hid herself in the cave forming a cover with the stones she was playing with. The stone that went up became the roof and the one fallen down became her seat.
The Mandir is on a hilltop which can be reached climbing 64 steps, inside a vast forest of 28 acres. It is protected by the ASI. It is believed that earlier Jain monks used the place for tapasya due to the presence of stone carved murtis of Jain Tirthankaras here.
The procession of Devi during the festival is on a female elephant. It is believed that once a tusker was used. Devi cursed the animal and turned it into a rock. The rock in the nearby land is still called elephant rock and is believed to be the elephant which was once cursed.

More from Jaya_Upadhyaya

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like