i go to huggingface, look on spaces just to see whats new and i see that MY OWN space is trending. ProCreations/maple-webgpu very cool to see, on the very first page of spaces too. anything is possible in the open source community!
I don't know if it was us or one of you guys or maybe all of us at once but lately we have seen a finetuning/pretraining explosion of models below 200m params and we can't be more happy about it keep coming tinkerers all of this is possible because of you!
Built a CPU-only BPM + musical key detector with Essentia and Gradio.
It returns BPM, key/scale, Camelot code, confidence, and analyzed duration. A representative 90-second window keeps CPU latency bounded; half-time/double-time and key-changing tracks remain the main edge cases.
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th. Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th. Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!