Row 6365
Content Data
This page contains data entry 6365 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
[**Vision Transformers Need Registers**](https://openreview.net/forum?id=2dnO3LLiJ1) *Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski*
**Abstract:** Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing.
[**Generalization in diffusion models arises from geometry-adaptive harmonic representations**](https://openreview.net/forum?id=ANvmVS2Yr0) *Zahra Kadkhodaie, Florentin Guth, Eero P Simoncelli, Stéphane Mallat*
**Abstract:** Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but recent reports of memorization of the training set raise the question of whether these networks are learning the “true” continuous density of the data. Here, we show that two DNNs trained on non-overlapping subsets of a dataset learn nearly the same score function, and thus the same density, when the number of training images is large enough. In this regime of strong generalization, diffusion-generated images are distinct from the training set, and are of high visual quality, suggesting that the inductive biases of the DNNs are well-aligned with the data density. We analyze the learned denoising functions and show that the inductive biases give rise to a shrinkage operation in a basis adapted to the underlying image. Examination of these bases reveals oscillating harmonic structures along contours and in homogeneous regions. We demonstrate that trained denoisers are inductively biased towards these geometry-adaptive harmonic bases since they arise not only when the network is trained on photographic images, but also when it is trained on image classes supported on low-dimensional manifolds for which the harmonic basis is suboptimal. Finally, we show that when trained on regular image classes for which the optimal basis is known to be geometry-adaptive and harmonic, the denoising performance of the networks is near-optimal.
[**Learning Interactive Real-World Simulators**](https://openreview.net/forum?id=sFyTZEqmUY) *Sherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, Pieter Abbeel*
**Abstract:** Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans, robots, and other interactive agents. Applications of a real-world simulator range from controllable content creation in games and movies, to training embodied agents purely in simulation that can be directly deployed in the real world. We explore the possibility of learning a universal simulator (UniSim) of real-world interaction through generative modeling. We first make the important observation that natural datasets available for learning a real-world simulator are often rich along different axes (e.g., abundant objects in image data, densely sampled actions in robotics data, and diverse movements in navigation data). With careful orchestration of diverse datasets, each providing a different aspect of the overall experience, UniSim can emulate how humans and agents interact with the world by simulating the visual outcome of both high-level instructions such as “open the drawer” and low-level controls such as “move by x,y” from otherwise static scenes and objects. There are numerous use cases for such a real-world simulator. As an example, we use UniSim to train both high-level vision-language planners and low-level reinforcement learning policies, each of which exhibit zero-shot real-world transfer after training purely in a learned real-world simulator. We also show that other types of intelligence such as video captioning models can benefit from training with simulated experience in UniSim, opening up even wider applications.
[**Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors**](https://openreview.net/forum?id=PdaPky8MUn) *Ido Amos, Jonathan Berant, Ankit Gupta*
**Abstract:** Modeling long-range dependencies across sequences is a longstanding goal in machine learning and has led to architectures, such as state space models, that dramatically outperform Transformers on long sequences. However, these impressive empirical gains have been by and large demonstrated on benchmarks (e.g. Long Range Arena), where models are randomly initialized and trained to predict a target label from an input sequence. In this work, we show that random initialization leads to gross overestimation of the differences between architectures and that pretraining with standard denoising objectives, using only the downstream task data, leads to dramatic gains across multiple architectures and to very small gaps between Transformers and state space models (SSMs). In stark contrast to prior works, we find vanilla Transformers to match the performance of S4 on Long Range Arena when properly pretrained, and we improve the best reported results of SSMs on the PathX-256 task by 20 absolute points. Subsequently, we analyze the utility of previously-proposed structured parameterizations for SSMs and show they become mostly redundant in the presence of data-driven initialization obtained through pretraining. Our work shows that, when evaluating different architectures on supervised tasks, incorporation of data-driven priors via pretraining is essential for reliable performance estimation, and can be done efficiently.
| Field | Value |
|---|---|
| text | [**Vision Transformers Need Registers**](https://openreview.net/forum?id=2dnO3LLiJ1) *Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski* **Abstract:** Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas o… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-09 |
| username_encoded | Z0FBQUFBQm5LakwyQk5CRFVoNl9PN09hcVM3N3RWSGxXU3BfSjhCR2tjUjhRZzhpMFdCLUh1ZDk3SFNKN1JiM2FRczViQkQ5eXZkZnhnS1lHRDhnNnZTSnpZYmdaNEVyU2c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HeW93SkFmX0JiQjd0OGY4OHdEQXVkWlo3TF9fcERtU01VYnB2YzZ5WW8tYzRZUUprbVFHd1hhNFptUUZNaktlaXdBMVFVNjZ5dXNOQmprUllMLWpBRmVLQmU2RHg2RkVqZXl5Tk9Dd1RJY2lkVzk3TkNsU3hjNlowZmpNS2JEYlpOS3k3d0lkYVY0SkxudFZWT3JKYm1Zb0dPWEFoVWQ0MHdSMjFnLTMyU3ViZi1KMmx3S19TTlFYcFhHLU51RmU5NDhQODFvV2dtaFJoTFN5dDlhQ0xEQT09 |
Raw Record
{
"text": "[**Vision Transformers Need Registers**](https://openreview.net/forum?id=2dnO3LLiJ1) \n*Timothée Darcet, Maxime Oquab, Julien Mairal, Piotr Bojanowski*\n\n**Abstract:** Transformers have recently emerged as a powerful tool for learning visual representations. In this paper, we identify and characterize artifacts in feature maps of both supervised and self-supervised ViT networks. The artifacts correspond to high-norm tokens appearing during inference primarily in low-informative background areas of images, that are repurposed for internal computations. We propose a simple yet effective solution based on providing additional tokens to the input sequence of the Vision Transformer to fill that role. We show that this solution fixes that problem entirely for both supervised and self-supervised models, sets a new state of the art for self-supervised visual models on dense visual prediction tasks, enables object discovery methods with larger models, and most importantly leads to smoother feature maps and attention maps for downstream visual processing.\n\n[**Generalization in diffusion models arises from geometry-adaptive harmonic representations**](https://openreview.net/forum?id=ANvmVS2Yr0) \n*Zahra Kadkhodaie, Florentin Guth, Eero P Simoncelli, Stéphane Mallat*\n\n**Abstract:** Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but recent reports of memorization of the training set raise the question of whether these networks are learning the “true” continuous density of the data. Here, we show that two DNNs trained on non-overlapping subsets of a dataset learn nearly the same score function, and thus the same density, when the number of training images is large enough. In this regime of strong generalization, diffusion-generated images are distinct from the training set, and are of high visual quality, suggesting that the inductive biases of the DNNs are well-aligned with the data density. We analyze the learned denoising functions and show that the inductive biases give rise to a shrinkage operation in a basis adapted to the underlying image. Examination of these bases reveals oscillating harmonic structures along contours and in homogeneous regions. We demonstrate that trained denoisers are inductively biased towards these geometry-adaptive harmonic bases since they arise not only when the network is trained on photographic images, but also when it is trained on image classes supported on low-dimensional manifolds for which the harmonic basis is suboptimal. Finally, we show that when trained on regular image classes for which the optimal basis is known to be geometry-adaptive and harmonic, the denoising performance of the networks is near-optimal.\n\n[**Learning Interactive Real-World Simulators**](https://openreview.net/forum?id=sFyTZEqmUY) \n*Sherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, Pieter Abbeel*\n\n**Abstract:** Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans, robots, and other interactive agents. Applications of a real-world simulator range from controllable content creation in games and movies, to training embodied agents purely in simulation that can be directly deployed in the real world. We explore the possibility of learning a universal simulator (UniSim) of real-world interaction through generative modeling. We first make the important observation that natural datasets available for learning a real-world simulator are often rich along different axes (e.g., abundant objects in image data, densely sampled actions in robotics data, and diverse movements in navigation data). With careful orchestration of diverse datasets, each providing a different aspect of the overall experience, UniSim can emulate how humans and agents interact with the world by simulating the visual outcome of both high-level instructions such as “open the drawer” and low-level controls such as “move by x,y” from otherwise static scenes and objects. There are numerous use cases for such a real-world simulator. As an example, we use UniSim to train both high-level vision-language planners and low-level reinforcement learning policies, each of which exhibit zero-shot real-world transfer after training purely in a learned real-world simulator. We also show that other types of intelligence such as video captioning models can benefit from training with simulated experience in UniSim, opening up even wider applications.\n\n[**Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors**](https://openreview.net/forum?id=PdaPky8MUn) \n*Ido Amos, Jonathan Berant, Ankit Gupta*\n\n**Abstract:** Modeling long-range dependencies across sequences is a longstanding goal in machine learning and has led to architectures, such as state space models, that dramatically outperform Transformers on long sequences. However, these impressive empirical gains have been by and large demonstrated on benchmarks (e.g. Long Range Arena), where models are randomly initialized and trained to predict a target label from an input sequence. In this work, we show that random initialization leads to gross overestimation of the differences between architectures and that pretraining with standard denoising objectives, using only the downstream task data, leads to dramatic gains across multiple architectures and to very small gaps between Transformers and state space models (SSMs). In stark contrast to prior works, we find vanilla Transformers to match the performance of S4 on Long Range Arena when properly pretrained, and we improve the best reported results of SSMs on the PathX-256 task by 20 absolute points. Subsequently, we analyze the utility of previously-proposed structured parameterizations for SSMs and show they become mostly redundant in the presence of data-driven initialization obtained through pretraining. Our work shows that, when evaluating different architectures on supervised tasks, incorporation of data-driven priors via pretraining is essential for reliable performance estimation, and can be done efficiently.",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-09",
"username_encoded": "Z0FBQUFBQm5LakwyQk5CRFVoNl9PN09hcVM3N3RWSGxXU3BfSjhCR2tjUjhRZzhpMFdCLUh1ZDk3SFNKN1JiM2FRczViQkQ5eXZkZnhnS1lHRDhnNnZTSnpZYmdaNEVyU2c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HeW93SkFmX0JiQjd0OGY4OHdEQXVkWlo3TF9fcERtU01VYnB2YzZ5WW8tYzRZUUprbVFHd1hhNFptUUZNaktlaXdBMVFVNjZ5dXNOQmprUllMLWpBRmVLQmU2RHg2RkVqZXl5Tk9Dd1RJY2lkVzk3TkNsU3hjNlowZmpNS2JEYlpOS3k3d0lkYVY0SkxudFZWT3JKYm1Zb0dPWEFoVWQ0MHdSMjFnLTMyU3ViZi1KMmx3S19TTlFYcFhHLU51RmU5NDhQODFvV2dtaFJoTFN5dDlhQ0xEQT09"
}
Entry Information
- Entry ID: 6365
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000