Collections
Discover the best community collections!
Collections including paper arxiv:2111.15664 
						
					
				- 
	
	
	
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Paper • 2306.17107 • Published • 11 - 
	
	
	
On the Hidden Mystery of OCR in Large Multimodal Models
Paper • 2305.07895 • Published • 1 - 
	
	
	
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Paper • 2308.12966 • Published • 11 - 
	
	
	
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Paper • 2401.15947 • Published • 53 
- 
	
	
	
UI Layout Generation with LLMs Guided by UI Grammar
Paper • 2310.15455 • Published • 3 - 
	
	
	
You Only Look at Screens: Multimodal Chain-of-Action Agents
Paper • 2309.11436 • Published • 1 - 
	
	
	
Never-ending Learning of User Interfaces
Paper • 2308.08726 • Published • 2 - 
	
	
	
LMDX: Language Model-based Document Information Extraction and Localization
Paper • 2309.10952 • Published • 66 
- 
	
	
	
DSG: An End-to-End Document Structure Generator
Paper • 2310.09118 • Published • 2 - 
	
	
	
OCR-free Document Understanding Transformer
Paper • 2111.15664 • Published • 4 - 
	
	
	
DocParser: End-to-end OCR-free Information Extraction from Visually Rich Documents
Paper • 2304.12484 • Published • 1 - 
	
	
	
Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration
Paper • 2309.01131 • Published • 1 
- 
	
	
	
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Paper • 2306.17107 • Published • 11 - 
	
	
	
On the Hidden Mystery of OCR in Large Multimodal Models
Paper • 2305.07895 • Published • 1 - 
	
	
	
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Paper • 2308.12966 • Published • 11 - 
	
	
	
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Paper • 2401.15947 • Published • 53 
- 
	
	
	
DSG: An End-to-End Document Structure Generator
Paper • 2310.09118 • Published • 2 - 
	
	
	
OCR-free Document Understanding Transformer
Paper • 2111.15664 • Published • 4 - 
	
	
	
DocParser: End-to-end OCR-free Information Extraction from Visually Rich Documents
Paper • 2304.12484 • Published • 1 - 
	
	
	
Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration
Paper • 2309.01131 • Published • 1 
- 
	
	
	
UI Layout Generation with LLMs Guided by UI Grammar
Paper • 2310.15455 • Published • 3 - 
	
	
	
You Only Look at Screens: Multimodal Chain-of-Action Agents
Paper • 2309.11436 • Published • 1 - 
	
	
	
Never-ending Learning of User Interfaces
Paper • 2308.08726 • Published • 2 - 
	
	
	
LMDX: Language Model-based Document Information Extraction and Localization
Paper • 2309.10952 • Published • 66