Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
Sperminator's profile picture
dlebedev's profile picture
Quazim0t0's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
replied
to
their
post
43 minutes ago
Say hello to the Trigram ByteLLM - AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 24000 steps aka 7b tokens in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
updated
a model
about 3 hours ago
AbstractPhil/alephllm-mini-beatrix-training
replied
to
their
post
about 21 hours ago
Say hello to the Trigram ByteLLM - AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 24000 steps aka 7b tokens in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say WILL be recorded in a public cache, WHEN the chat version works. For now she records nothing. The idea is to help debug the K/V cache, and I would rather the data accumulated be shared. If you wish to speak to her in private I will include a toggle, that way you'll see that nothing is recorded when you speak and you can still have a private chat with her. For now she's simply auto-completing, so have fun with her. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
View all activity
Organizations
AbstractPhil
's models
213
Sort: Recently updated
AbstractPhil/alephllm-mini-beatrix-training
Updated
19 minutes ago
AbstractPhil/aleph-splat-0
Updated
2 days ago
AbstractPhil/alephlm-0
Feature Extraction
•
Updated
5 days ago
AbstractPhil/alephlm-adopt-0
Text Generation
•
Updated
5 days ago
AbstractPhil/captionbert-8192-v2-b
Feature Extraction
•
58.3M
•
Updated
5 days ago
•
93
•
1
AbstractPhil/captionbert-8192-v2
Feature Extraction
•
58.3M
•
Updated
5 days ago
•
227
•
1
AbstractPhil/sd15-flow-lune
Text-to-Image
•
Updated
10 days ago
•
18
AbstractPhil/clip-vitb-mini-distilled
Image Feature Extraction
•
8.93M
•
Updated
11 days ago
•
408
AbstractPhil/loss-manifest
Updated
11 days ago
AbstractPhil/geolip-bertenstein
Feature Extraction
•
Updated
13 days ago
AbstractPhil/geolip-vit-captionbank-coco
Image Feature Extraction
•
Updated
16 days ago
AbstractPhil/geolip-vit-base-x3
11.7M
•
Updated
17 days ago
•
48
AbstractPhil/geolip-vit-large-x3
78.3M
•
Updated
17 days ago
•
31
AbstractPhil/geolip-aleph-diffusion
Updated
19 days ago
•
2
AbstractPhil/geolip-aleph-qwen-3.5-0.8b-instruct
Updated
19 days ago
•
1
AbstractPhil/amoe-lora
Updated
23 days ago
AbstractPhil/aleph-diffusion-adapters
Updated
25 days ago
AbstractPhil/qwen3.5-0.8b-relay-caption
Updated
27 days ago
AbstractPhil/geolip-aleph-qwen
Updated
29 days ago
AbstractPhil/geolip-aleph-differentiation
Updated
Jul 12
AbstractPhil/anima-90k
Updated
Jul 7
•
1
AbstractPhil/geolip-aleph-lm
Text Generation
•
Updated
Jul 1
•
1
AbstractPhil/qwen-benchmark
Updated
Jun 30
AbstractPhil/anima-brent-10k
Updated
Jun 28
AbstractPhil/anima-prelim-1k-r64
Text-to-Image
•
Updated
Jun 25
•
1
AbstractPhil/Qwen3.5-0.8B-json-captioner
Image-Text-to-Text
•
0.9B
•
Updated
Jun 25
•
58
AbstractPhil/geolip-constellation-aleph
Updated
Jun 19
AbstractPhil/geolip-aleph-void
Feature Extraction
•
Updated
Jun 14
AbstractPhil/geolip-sdxl-aleph
Text-to-Image
•
Updated
Jun 8
•
•
2
AbstractPhil/geolip-hypersphere-experiments
Updated
Jun 3
•
1
Previous
1
2
3
...
8
Next