Row 91078

Row ID: 91078 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 91078 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

V-JEPA idea is cool and all, but I don’t see any subsequent works after it. I have tried doing a PCA projection on the features extracted from the encoder and visualize them. What makes me stumbled was that the initial weight of the backbone captured the structure of the clips better than the pre-trained V-JEPA (I used Nvidia’s RADIO example code for it)

Does anyone have similar experience that they could share with.

Btw, I posted an issue on V-JEPA Github. You could see the feature visualization there in the issue and we could discuss more technical details there. I just think that people might be more active here in the community.

https://github.com/facebookresearch/jepa/issues/66

FieldValue
text V-JEPA idea is cool and all, but I don’t see any subsequent works after it. I have tried doing a PCA projection on the features extracted from the encoder and visualize them. What makes me stumbled was that the initial weight of the backbone captured the structure of the clips better than the pre-trained V-JEPA (I used Nvidia’s RADIO example code for it) Does anyone have similar experience that they could share with. Btw, I posted an issue on V-JEPA Github. You could see the feature visualizat…
label r/deeplearning
dataType post
communityName r/deeplearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak1yOHRNWVoxZkZIVDItQU9hdFVYMlV1WFJ0Ynp0bDNYRklTNVBfN0h4MkU0TDZDT1BCeW5Ga3ptd1RVc2J2Y3IxQmE1T2U5eVZZMFJyRlJNamI0MDc3M0E9PQ==
url_encoded Z0FBQUFBQm5Lak85MHVzd0w5VWluRjItb2FNSWJvemJTTkRxN0QwUHhISm80MHFVRnBfNGMxWE9PaHdhYkhvRnQxRzh3MG9wTTJ2MFBlYzNlVGhGZXdVSXA0U2phRkRlTlZEOGhWSXo2ajhwTDFWNVlZZzh6VnVXZ0FoWDdZTDhYRUZvS2RRNy14LTlTOEtHdUxBSy1CMjZDSE5GSzhqUE1DTWhJYVI2SXladFB1Y09IT2Q4cm9kYkl2OUh5UUZEakRpNFYyeEtpenVu

Raw Record

{
  "text": "V-JEPA idea is cool and all, but I don’t see any subsequent works after it. I have tried doing a PCA projection on the features extracted from the encoder and visualize them. What makes me stumbled was that the initial weight of the backbone captured the structure of the clips better than the pre-trained V-JEPA (I used Nvidia’s RADIO example code for it)\n\nDoes anyone have similar experience that they could share with.\n\nBtw, I posted an issue on V-JEPA Github. You could see the feature visualization there in the issue and we could discuss more technical details there. I just think that people might be more active here in the community.\n\n\nhttps://github.com/facebookresearch/jepa/issues/66",
  "label": "r/deeplearning",
  "dataType": "post",
  "communityName": "r/deeplearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak1yOHRNWVoxZkZIVDItQU9hdFVYMlV1WFJ0Ynp0bDNYRklTNVBfN0h4MkU0TDZDT1BCeW5Ga3ptd1RVc2J2Y3IxQmE1T2U5eVZZMFJyRlJNamI0MDc3M0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak85MHVzd0w5VWluRjItb2FNSWJvemJTTkRxN0QwUHhISm80MHFVRnBfNGMxWE9PaHdhYkhvRnQxRzh3MG9wTTJ2MFBlYzNlVGhGZXdVSXA0U2phRkRlTlZEOGhWSXo2ajhwTDFWNVlZZzh6VnVXZ0FoWDdZTDhYRUZvS2RRNy14LTlTOEtHdUxBSy1CMjZDSE5GSzhqUE1DTWhJYVI2SXladFB1Y09IT2Q4cm9kYkl2OUh5UUZEakRpNFYyeEtpenVu"
}

Entry Information