Stanislaw Jastrzebski

A Closer Look at Memorization in Deep Networks

Devansh Arpit

Stanisław Jastrzębski

Maxinder S. Kanwal

We examine the role of memorization in deep learning, drawing connections to capacity, generalization, and adversarial robustness. While dee… (voir plus)p networks are capable of memorizing noise data, our results suggest that they tend to prioritize learning simple patterns first. In our experiments, we expose qualitative differences in gradient-based optimization of deep neural networks (DNNs) on noise vs. real data. We also demonstrate that for appropriately tuned explicit regularization (e.g., dropout) we can degrade DNN training performance on noise datasets without compromising generalization on real data. Our analysis suggests that the notions of effective capacity which are dataset independent are unlikely to explain the generalization performance of deep networks when trained with gradient based methods because training data itself plays an important role in determining the degree of memorization.

2017-07-17

Proceedings of the 34th International Conference on Machine Learning (publié)

proceedings.mlr.press

arxiv.org

A Closer Look at Memorization in Deep Networks

Devansh Arpit

Stanisław Jastrzębski

Maxinder S. Kanwal

We examine the role of memorization in deep learning, drawing connections to capacity, generalization, and adversarial robustness. While dee… (voir plus)p networks are capable of memorizing noise data, our results suggest that they tend to prioritize learning simple patterns first. In our experiments, we expose qualitative differences in gradient-based optimization of deep neural networks (DNNs) on noise vs. real data. We also demonstrate that for appropriately tuned explicit regularization (e.g., dropout) we can degrade DNN training performance on noise datasets without compromising generalization on real data. Our analysis suggests that the notions of effective capacity which are dataset independent are unlikely to explain the generalization performance of deep networks when trained with gradient based methods because training data itself plays an important role in determining the degree of memorization.

2017-06-16

ArXiv (prépublication)

arxiv.org

Deep Nets Don't Learn via Memorization

David Scott Krueger

Nicolas Ballas

Stanisław Jastrzębski

Maxinder S. Kanwal

We use empirical methods to argue that deep neural networks (DNNs) do not achieve their performance by memorizing training data in spite of … (voir plus)overlyexpressive model architectures. Instead, they learn a simple available hypothesis that fits the finite data samples. In support of this view, we establish that there are qualitative differences when learning noise vs. natural datasets, showing: (1) more capacity is needed to fit noise, (2) time to convergence is longer for random labels, but shorter for random inputs, and (3) that DNNs trained on real data examples learn simpler functions than when trained with noise data, as measured by the sharpness of the loss function at convergence. Finally, we demonstrate that for appropriately tuned explicit regularization, e.g. dropout, we can degrade DNN training performance on noise datasets without compromising generalization on real data.

2017-02-17

International Conference on Learning Representations (publié)

dblp.uni-trier.de

À l’avant-garde d’une nouvelle ère

Éclaireurs autochtones en IA

Avantage IA

Stanislaw Jastrzebski

Publications

À l’avant-garde d’une nouvelle ère

Éclaireurs autochtones en IA

Avantage IA

Mots-clés populaires:

Stanislaw Jastrzebski

Publications