Hey everyone! I got sidetracked from my main projects and decided to test out the BananaAll app and see if I could make a small model not regress too much during SFT. Here is what happened:
- Thank you to @Banaxi-Tech for the BananaAll app (works perfectly on Windows and CPU) - GPT-6 Sol for knowing how to merge some confusing files created by the app - Myself for the idea - Someone else somewhere who might have contributed to some of my ideas and might in the future - And readers like you!
We're working on Pebble 1.5. Here's what we know so far:
- They will be better than the last generation. 99.99% certain. - Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768. - A Mamba3-based architecture with some other new architectural designs we're experimenting with. - Native CPU compatibility — something we failed at with the last generation. - Natively multilingual and multimodal???
2. SmolCodeBench
A code benchmark designed specifically for small models, because there really isn't a good one right now.
3. SENTRY
VOID is working on something called SENTRY — System for Evaluating Neural Threats, Responses, and Yields.
More on that soon.
4. basically OS
It's an operating system/app/harness. We're still deciding.
5. Finances
Trying to balance the finances after purchasing a Hugging Face Pro subscription.
G1-MINI has now seen around 8B tokens during its current run, and pretraining is still going strong.
Our E1 (Efficiency-1) prototype has also reached 15B pretraining tokens. E1 has 1B total parameters while activating under 100M parameters per token. It features adaptive activation, meaning easier tokens can use less compute while harder tokens receive more.
We plan to open-source E1 ASAP! 🚀
We’re also excited to announce Project Prism, which will provide limited access to our upcoming Orion Flagship model, powered by our T2 architecture.
Note: T2 here refers to the architecture, not our T2 (Thinker-2) model.
Applications for Project Prism are available through the org page, with more details coming soon!
Finally, welcome @soyL061215, who joined the Hugging Face org today! 🎉
Hello Everyone! I am happy to announce a few things. 1. G1-MINI G1-MINI is now in pretraining and is training at a steady pace. Our current ETAs state completion and launch in about 15-20 days, somewhere near the end of September. 2. G1-NANO G1-NANO is also being pretrained as we speak at a pace of over 400K tokens per second processing more than 10B tokens in 12 hours. This allows us to train extremely fast, and we will launch it somewhere around 15th September. 3. We have begun work on FrameShot, a dual image and video generation model at around 4B dense parameters. This is expected to launch around late December with no promised date. - Bc-AI
Hello everyone! I have 2 announcements today! The first one is the launch of our new API platform! You can make a account and get 5 dollars free credits. No credits card needed because i have no idea how to set up a payment's thing. If you want more credits just email me at smilyai@outlook.com . The platform currently features G1-Preview a preview of G1 and the older Mira-1-Large. 2nd announcement is we have started working on G1-MINI so expect a late October Ish launch - Bc-AI on behalf of Smilyai-Labs
There are still some interesting improvements in these models: - Compatible with non-CUDA devices - Vocabulary increased to 16K tokens - Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
I know a lot of you were really looking forward to the original 20B MoE, and honestly, I was really excited about it too.
Unfortunately, the free compute credits I was using from ML Intern Explorers were removed by Hugging Face. That changed what I can realistically do with the original plan, so I've decided to move G1 over to a Qwen 3.8 27B base instead. Its not just another finetune though, I am inserting extra layers and putting it through my vigourous pipeline. Results will be open source.
I know that's probably disappointing, especially for the people who were specifically waiting for the 20B MoE. I'm genuinely sorry about that.
I really appreciate everyone who got excited about G1 in the first place. I didn't expect this change either, but I'm still really excited to see where G1 can go from here.
The original G1 codebase will stay open too. Its under my profile: Bc-AI/train-g1
I have some unfortunate news to share with everyone. My earlier estimate for the launch of G1 in late October was inaccurate. We sincerely apologise for any inconvenience this may cause, but with our current compute resources, pretraining a 20B MoE model is not realistically possible within two months.
G1-MINI will also be postponed, but not for nearly as long — only by a few extra months.
This is disappointing, as I know I was excited to launch G1, and I know many people were also watching the model and looking forward to it.
However in my view, I would rather be honest about our limitations than be overly optimistic about something we currently cannot guarantee.
G1 is not cancelled but uh it will be postponed indefinitely until I have the resources needed to train it. This could be next month, or it could take years. For now, I don't want to give another estimated launch date until I know we have the resources to actually make it happen.
In the meantime, SmilyAI will continue developing AI and experimenting with new ideas, and we will provide updates as we go.
Thank you for sticking with us and supporting SmilyAI. We will continue working towards better models in the future.
Also, if you do have the hardware, to run it aka 8xH200s or better, my codebase is fully open under my very permissive license: i-have-no-idea-just-use-this. Basically, do whatever just mention me. Bc-AI/train-g1
Hello everyone! Me and the team are working on G1-MINI and G1. Right now, G1-MINI is aimed at a launch in mid to late September, depending on how fast we fix the minor issues. As for G1, it's looking like a late October to mid-November launch, based on current trajectory. If things go terribly wrong, we could postpone it to December, as we prefer to ship confidently, not ship a half-done dogs' breakfast of a model. 🤣 All dates could be changed at any moment, as we are high school students not full-time ML engineers 😅. Other things to look out for is an overhaul of the UI and the information on my website. Thanks to my beta testers: @guardamarcos@Timmy6767@MUK-IS-GOAT@smilyai-large-team@Sbui503@Banaxi-Tech@Bc-AI@atom77777@Harley-ml@Datdanboi25@Fishtiks@smartdigitalnetworks@vovaRL@EmetTheGolum@juiceb0xc0de@ProCreations