-
Attention Is All You Need
Paper β’ 1706.03762 β’ Published β’ 140 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper β’ 1912.01703 β’ Published β’ 2 -
google-bert/bert-base-uncased
Fill-Mask β’ 0.1B β’ Updated β’ 69.7M β’ β’ 2.83k -
openai-community/gpt2
Text Generation β’ 0.1B β’ Updated β’ 14.5M β’ 3.51k
Collections
Discover the best community collections!
Collections including paper arxiv:2103.00020
-
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Paper β’ 2511.22699 β’ Published β’ 249 -
A Survey on Diffusion Language Models
Paper β’ 2508.10875 β’ Published β’ 35 -
High-Resolution Image Synthesis with Latent Diffusion Models
Paper β’ 2112.10752 β’ Published β’ 17 -
Denoising Diffusion Probabilistic Models
Paper β’ 2006.11239 β’ Published β’ 9
-
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
Paper β’ 2010.14406 β’ Published -
Learning Transferable Visual Models From Natural Language Supervision
Paper β’ 2103.00020 β’ Published β’ 22 -
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Paper β’ 2010.11929 β’ Published β’ 23
-
Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning
Paper β’ 2211.04325 β’ Published β’ 1 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper β’ 1810.04805 β’ Published β’ 33 -
On the Opportunities and Risks of Foundation Models
Paper β’ 2108.07258 β’ Published β’ 2 -
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Paper β’ 2204.07705 β’ Published β’ 2
-
Neural Machine Translation by Jointly Learning to Align and Translate
Paper β’ 1409.0473 β’ Published β’ 7 -
Attention Is All You Need
Paper β’ 1706.03762 β’ Published β’ 140 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper β’ 1810.04805 β’ Published β’ 33 -
Hierarchical Reasoning Model
Paper β’ 2506.21734 β’ Published β’ 54
-
sentence-transformers/all-mpnet-base-v2
Sentence Similarity β’ 0.1B β’ Updated β’ 24.6M β’ β’ 1.35k -
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Paper β’ 1910.10683 β’ Published β’ 20 -
google-t5/t5-base
Translation β’ 0.2B β’ Updated β’ 3.94M β’ β’ 790 -
Attention Is All You Need
Paper β’ 1706.03762 β’ Published β’ 140
-
MIO: A Foundation Model on Multimodal Tokens
Paper β’ 2409.17692 β’ Published β’ 53 -
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Paper β’ 2010.11929 β’ Published β’ 23 -
Going deeper with Image Transformers
Paper β’ 2103.17239 β’ Published -
Training data-efficient image transformers & distillation through attention
Paper β’ 2012.12877 β’ Published β’ 2
-
Attention Is All You Need
Paper β’ 1706.03762 β’ Published β’ 140 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper β’ 1912.01703 β’ Published β’ 2 -
google-bert/bert-base-uncased
Fill-Mask β’ 0.1B β’ Updated β’ 69.7M β’ β’ 2.83k -
openai-community/gpt2
Text Generation β’ 0.1B β’ Updated β’ 14.5M β’ 3.51k
-
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Paper β’ 2511.22699 β’ Published β’ 249 -
A Survey on Diffusion Language Models
Paper β’ 2508.10875 β’ Published β’ 35 -
High-Resolution Image Synthesis with Latent Diffusion Models
Paper β’ 2112.10752 β’ Published β’ 17 -
Denoising Diffusion Probabilistic Models
Paper β’ 2006.11239 β’ Published β’ 9
-
Neural Machine Translation by Jointly Learning to Align and Translate
Paper β’ 1409.0473 β’ Published β’ 7 -
Attention Is All You Need
Paper β’ 1706.03762 β’ Published β’ 140 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper β’ 1810.04805 β’ Published β’ 33 -
Hierarchical Reasoning Model
Paper β’ 2506.21734 β’ Published β’ 54
-
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
Paper β’ 2010.14406 β’ Published -
Learning Transferable Visual Models From Natural Language Supervision
Paper β’ 2103.00020 β’ Published β’ 22 -
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Paper β’ 2010.11929 β’ Published β’ 23
-
Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning
Paper β’ 2211.04325 β’ Published β’ 1 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper β’ 1810.04805 β’ Published β’ 33 -
On the Opportunities and Risks of Foundation Models
Paper β’ 2108.07258 β’ Published β’ 2 -
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Paper β’ 2204.07705 β’ Published β’ 2
-
sentence-transformers/all-mpnet-base-v2
Sentence Similarity β’ 0.1B β’ Updated β’ 24.6M β’ β’ 1.35k -
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Paper β’ 1910.10683 β’ Published β’ 20 -
google-t5/t5-base
Translation β’ 0.2B β’ Updated β’ 3.94M β’ β’ 790 -
Attention Is All You Need
Paper β’ 1706.03762 β’ Published β’ 140
-
MIO: A Foundation Model on Multimodal Tokens
Paper β’ 2409.17692 β’ Published β’ 53 -
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Paper β’ 2010.11929 β’ Published β’ 23 -
Going deeper with Image Transformers
Paper β’ 2103.17239 β’ Published -
Training data-efficient image transformers & distillation through attention
Paper β’ 2012.12877 β’ Published β’ 2