Simple Self-Distillation
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
Environment-free Synthetic Data Generation for API-Calling Agents
TopoPrimer: The Missing Topological Context in Forecasting Models
Team members 736 private
Efficient Vision Encoding for Vision Language Models
-
FastVLM: Efficient Vision Encoding for Vision Language Models
Paper β’ 2412.13303 β’ Published β’ 77 -
FastVLM WebGPU
π446Real-time video captioning powered by FastVLM
-
apple/FastVLM-0.5B
Text Generation β’ 0.8B β’ Updated β’ 20.5k β’ 396 -
apple/FastVLM-1.5B
Text Generation β’ 2B β’ Updated β’ 7.32k β’ 80
-
apple/coreml-depth-anything-v2-small
Depth Estimation β’ Updated β’ 1.01k β’ 100 -
apple/coreml-depth-anything-small
Depth Estimation β’ Updated β’ 125 β’ 39 -
apple/coreml-detr-semantic-segmentation
Image Segmentation β’ Updated β’ 165 β’ 33 -
apple/coreml-FastViT-T8
Image Classification β’ Updated β’ 23 β’ 18
Benchmark for the design of efficient continual learning of image-text models over years.
-
TiC-CLIP: Continual Training of CLIP Models
Paper β’ 2310.16226 β’ Published β’ 10 -
apple/TiC-DataComp
Preview β’ Updated β’ 1.64k β’ 5 -
apple/TiC-CLIP-basic-cumulative
Zero-Shot Image Classification β’ Updated β’ 103 β’ 3 -
apple/TiC-CLIP-basic-oracle
Zero-Shot Image Classification β’ Updated β’ 8 β’ 1
-
apple/coreml-stable-diffusion-mixed-bit-palettization
Updated β’ 1 β’ 30 -
apple/coreml-stable-diffusion-xl-base
Text-to-Image β’ Updated β’ 50 β’ 70 -
apple/coreml-stable-diffusion-2-1-base
Text-to-Image β’ Updated β’ 44 β’ 56 -
pcuenq/coreml-stable-diffusion-2-1-base
Text-to-Image β’ Updated β’ 42 β’ 4
AIM: Autoregressive Image Models
CLaRa models
MobileCLIP2: Mobile-friendly image-text models with SOTA zero-shot capabilities trained on DFNDR-2B
A collection of AIMv2 vision encoders that supports a number of resolutions, native resolution, and a distilled checkpoint.
-
apple/aimv2-large-patch14-224
Image Feature Extraction β’ 0.3B β’ Updated β’ 2.86k β’ 62 -
apple/aimv2-huge-patch14-224
Image Feature Extraction β’ 0.7B β’ Updated β’ 270 β’ 13 -
apple/aimv2-1B-patch14-224
Image Feature Extraction β’ 1B β’ Updated β’ 167 β’ 8 -
apple/aimv2-3B-patch14-224
Image Feature Extraction β’ 3B β’ Updated β’ 76 β’ 4
-
apple/OpenELM-270M-Instruct
Text Generation β’ 0.3B β’ Updated β’ 1.26k β’ 145 -
apple/OpenELM-450M-Instruct
Text Generation β’ 0.5B β’ Updated β’ 897 β’ 50 -
apple/OpenELM-1_1B-Instruct
Text Generation β’ 1B β’ Updated β’ 1.42M β’ 75 -
apple/OpenELM-3B-Instruct
Text Generation β’ 3B β’ Updated β’ 433 β’ 338
MobileCLIP: Mobile-friendly image-text models with SOTA zero-shot capabilities.
DataCompDR: Improved datasets for training image-text SOTA models.
-
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
Paper β’ 2311.17049 β’ Published β’ 8 -
apple/mobileclip_s0_timm
Image Classification β’ Updated β’ 1.02k β’ 12 -
apple/mobileclip_s1_timm
Image Classification β’ Updated β’ 105 β’ 3 -
apple/mobileclip_s2_timm
Image Classification β’ Updated β’ 126 β’ 6
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
-
apple/DepthPro-hf
Depth Estimation β’ 1.0B β’ Updated β’ 29.6k β’ 108 -
apple/DepthPro
Depth Estimation β’ Updated β’ 9.85k β’ 527 -
apple/DepthPro-mixin
Depth Estimation β’ 1.0B β’ Updated β’ 21 β’ 8 -
openai/clip-vit-large-patch14
Zero-Shot Image Classification β’ 0.4B β’ Updated β’ 7.8M β’ 2.07k
CLIP Models trained using DFN-2B/DFN-5B datasets
DCLM Models + Datasets
Simple Self-Distillation
CLaRa models
Efficient Vision Encoding for Vision Language Models
-
FastVLM: Efficient Vision Encoding for Vision Language Models
Paper β’ 2412.13303 β’ Published β’ 77 -
FastVLM WebGPU
π446Real-time video captioning powered by FastVLM
-
apple/FastVLM-0.5B
Text Generation β’ 0.8B β’ Updated β’ 20.5k β’ 396 -
apple/FastVLM-1.5B
Text Generation β’ 2B β’ Updated β’ 7.32k β’ 80
MobileCLIP2: Mobile-friendly image-text models with SOTA zero-shot capabilities trained on DFNDR-2B
A collection of AIMv2 vision encoders that supports a number of resolutions, native resolution, and a distilled checkpoint.
-
apple/aimv2-large-patch14-224
Image Feature Extraction β’ 0.3B β’ Updated β’ 2.86k β’ 62 -
apple/aimv2-huge-patch14-224
Image Feature Extraction β’ 0.7B β’ Updated β’ 270 β’ 13 -
apple/aimv2-1B-patch14-224
Image Feature Extraction β’ 1B β’ Updated β’ 167 β’ 8 -
apple/aimv2-3B-patch14-224
Image Feature Extraction β’ 3B β’ Updated β’ 76 β’ 4
-
apple/coreml-depth-anything-v2-small
Depth Estimation β’ Updated β’ 1.01k β’ 100 -
apple/coreml-depth-anything-small
Depth Estimation β’ Updated β’ 125 β’ 39 -
apple/coreml-detr-semantic-segmentation
Image Segmentation β’ Updated β’ 165 β’ 33 -
apple/coreml-FastViT-T8
Image Classification β’ Updated β’ 23 β’ 18
-
apple/OpenELM-270M-Instruct
Text Generation β’ 0.3B β’ Updated β’ 1.26k β’ 145 -
apple/OpenELM-450M-Instruct
Text Generation β’ 0.5B β’ Updated β’ 897 β’ 50 -
apple/OpenELM-1_1B-Instruct
Text Generation β’ 1B β’ Updated β’ 1.42M β’ 75 -
apple/OpenELM-3B-Instruct
Text Generation β’ 3B β’ Updated β’ 433 β’ 338
MobileCLIP: Mobile-friendly image-text models with SOTA zero-shot capabilities.
DataCompDR: Improved datasets for training image-text SOTA models.
-
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
Paper β’ 2311.17049 β’ Published β’ 8 -
apple/mobileclip_s0_timm
Image Classification β’ Updated β’ 1.02k β’ 12 -
apple/mobileclip_s1_timm
Image Classification β’ Updated β’ 105 β’ 3 -
apple/mobileclip_s2_timm
Image Classification β’ Updated β’ 126 β’ 6
Benchmark for the design of efficient continual learning of image-text models over years.
-
TiC-CLIP: Continual Training of CLIP Models
Paper β’ 2310.16226 β’ Published β’ 10 -
apple/TiC-DataComp
Preview β’ Updated β’ 1.64k β’ 5 -
apple/TiC-CLIP-basic-cumulative
Zero-Shot Image Classification β’ Updated β’ 103 β’ 3 -
apple/TiC-CLIP-basic-oracle
Zero-Shot Image Classification β’ Updated β’ 8 β’ 1
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
-
apple/DepthPro-hf
Depth Estimation β’ 1.0B β’ Updated β’ 29.6k β’ 108 -
apple/DepthPro
Depth Estimation β’ Updated β’ 9.85k β’ 527 -
apple/DepthPro-mixin
Depth Estimation β’ 1.0B β’ Updated β’ 21 β’ 8 -
openai/clip-vit-large-patch14
Zero-Shot Image Classification β’ 0.4B β’ Updated β’ 7.8M β’ 2.07k
-
apple/coreml-stable-diffusion-mixed-bit-palettization
Updated β’ 1 β’ 30 -
apple/coreml-stable-diffusion-xl-base
Text-to-Image β’ Updated β’ 50 β’ 70 -
apple/coreml-stable-diffusion-2-1-base
Text-to-Image β’ Updated β’ 44 β’ 56 -
pcuenq/coreml-stable-diffusion-2-1-base
Text-to-Image β’ Updated β’ 42 β’ 4
CLIP Models trained using DFN-2B/DFN-5B datasets
AIM: Autoregressive Image Models
DCLM Models + Datasets