-
google/pix2struct-screen2words-large
Visual Question Answering • Updated • 68 • 22 -
google/pix2struct-widget-captioning-large
Visual Question Answering • 1B • Updated • 32 • 20 -
google/matcha-chart2text-pew
Visual Question Answering • Updated • 33 • 40 -
jadechoghari/Ferret-UI-Gemma2b
Image-Text-to-Text • Updated • 832 • 52
Collections
Discover the best community collections!
Collections including paper arxiv:2310.09199
-
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Paper • 2305.06500 • Published • 6 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28 -
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Paper • 2306.05424 • Published • 7 -
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Paper • 2409.01704 • Published • 83
-
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28 -
A Zero-Shot Language Agent for Computer Control with Structured Reflection
Paper • 2310.08740 • Published • 15 -
Personality Traits in Large Language Models
Paper • 2307.00184 • Published • 21 -
An Emulator for Fine-Tuning Large Language Models using Small Language Models
Paper • 2310.12962 • Published • 13
-
Self-Alignment with Instruction Backtranslation
Paper • 2308.06259 • Published • 43 -
ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation
Paper • 2308.03793 • Published • 12 -
From Sparse to Soft Mixtures of Experts
Paper • 2308.00951 • Published • 22 -
Revisiting DETR Pre-training for Object Detection
Paper • 2308.01300 • Published • 10
-
Attention Is All You Need
Paper • 1706.03762 • Published • 120 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 26 -
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Paper • 1907.11692 • Published • 10 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 22
-
Table-GPT: Table-tuned GPT for Diverse Table Tasks
Paper • 2310.09263 • Published • 40 -
A Zero-Shot Language Agent for Computer Control with Structured Reflection
Paper • 2310.08740 • Published • 15 -
The Consensus Game: Language Model Generation via Equilibrium Search
Paper • 2310.09139 • Published • 14 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28
-
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
Paper • 2309.14717 • Published • 46 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28 -
Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams
Paper • 2310.08678 • Published • 13 -
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Paper • 2310.09478 • Published • 21
-
Vision Transformer Adapters for Generalizable Multitask Learning
Paper • 2308.12372 • Published -
RMT: Retentive Networks Meet Vision Transformers
Paper • 2309.11523 • Published • 34 -
DualToken-ViT: Position-aware Efficient Vision Transformer with Dual Token Fusion
Paper • 2309.12424 • Published • 11 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28
-
google/pix2struct-screen2words-large
Visual Question Answering • Updated • 68 • 22 -
google/pix2struct-widget-captioning-large
Visual Question Answering • 1B • Updated • 32 • 20 -
google/matcha-chart2text-pew
Visual Question Answering • Updated • 33 • 40 -
jadechoghari/Ferret-UI-Gemma2b
Image-Text-to-Text • Updated • 832 • 52
-
Attention Is All You Need
Paper • 1706.03762 • Published • 120 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 26 -
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Paper • 1907.11692 • Published • 10 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 22
-
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Paper • 2305.06500 • Published • 6 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28 -
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Paper • 2306.05424 • Published • 7 -
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Paper • 2409.01704 • Published • 83
-
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28 -
A Zero-Shot Language Agent for Computer Control with Structured Reflection
Paper • 2310.08740 • Published • 15 -
Personality Traits in Large Language Models
Paper • 2307.00184 • Published • 21 -
An Emulator for Fine-Tuning Large Language Models using Small Language Models
Paper • 2310.12962 • Published • 13
-
Table-GPT: Table-tuned GPT for Diverse Table Tasks
Paper • 2310.09263 • Published • 40 -
A Zero-Shot Language Agent for Computer Control with Structured Reflection
Paper • 2310.08740 • Published • 15 -
The Consensus Game: Language Model Generation via Equilibrium Search
Paper • 2310.09139 • Published • 14 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28
-
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
Paper • 2309.14717 • Published • 46 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28 -
Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams
Paper • 2310.08678 • Published • 13 -
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Paper • 2310.09478 • Published • 21
-
Self-Alignment with Instruction Backtranslation
Paper • 2308.06259 • Published • 43 -
ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation
Paper • 2308.03793 • Published • 12 -
From Sparse to Soft Mixtures of Experts
Paper • 2308.00951 • Published • 22 -
Revisiting DETR Pre-training for Object Detection
Paper • 2308.01300 • Published • 10
-
Vision Transformer Adapters for Generalizable Multitask Learning
Paper • 2308.12372 • Published -
RMT: Retentive Networks Meet Vision Transformers
Paper • 2309.11523 • Published • 34 -
DualToken-ViT: Position-aware Efficient Vision Transformer with Dual Token Fusion
Paper • 2309.12424 • Published • 11 -
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Paper • 2310.09199 • Published • 28