Imagine having some kind of server so it always feels updated with your latest integrated models!
appvoid
AI & ML interests
Recent Activity
Organizations
- Complete modern UI redesign
- Adds support for Qwen3.5 0.8B, LFM2.5 230M,350M, SmolLM2 360M, Gemma 3 270M.
- Adds Q7,Q6,Q5,Q3,Q1 quantization formats with a easy to use precision slider
- And more!
The new UI includes:
- New 1024ร768 High Quality interface.
- Photographic QOI background.
- Transparent BananaMind, CPU, cube, mouse, and Send icons.
- Proper bitmap cursor.
- Rounded translucent panels and cards.
- Modern model-loading progress window.
- Redesigned inference screen with response and prompt panels.
- Localized redraws for the cursor, clicks, loading progress, and precision slider.
Notice: Qwen3.5 0.8B currently generates garbled text, it will be fixed tomorrow.
See it for yourself
Now Available at https://github.com/BananaMind/BananaMindOS
Prebuild ISOs coming soon!
(also press ? + G if you want to load try to load a 6MB RAM model on 5MB may break)
byte-level tokenizer is the part that caught my attention, keep it up sir!
Let's go!!!
Yes sir!
Do you mean like the architecture? Could you point at the model?
Waiting for your new models sir
Hey @appvoid
Are you gonna press it?
Done. hahaha bots are becoming more of a thing here lately.
Hmm interesting!
You should only use some loops, like lets say you have 4 layers then:
L1 ๐ ฎ L2 ๐ ฎ L2 ๐ ฎ L3 ๐ ฎ L4 so 5 effective layers notL1 ๐ ฎ L1 ๐ ฎ L2 ๐ ฎ L2 ๐ ฎ L3 ๐ ฎ L3 ๐ ฎ L4 ๐ ฎ L4 because our tests on 1M:
Metric All looped (6 blocks) Partial looped (4 blocks) No loop (3 blocks) Base Bench Elo 875 885 884 Base raw accuracy 33.14% 34.57% 33.71% Base weighted accuracy 32.54% 33.65% 33.57% ARC Easy acc_norm 30.98% 29.50% 30.35% ARC Challenge acc_norm 21.33% 22.53% 22.10% PIQA acc_norm 52.45% 54.30% 52.56% HellaSwag acc_norm 27.28% 27.04% 26.93% ArithMark-3 acc_norm 30.40% 33.00% 30.80% INT Index 3.88 5.37 3.93 Training throughput 344K tok/s 492K tok/s 553K tok/s
It's always 4 steps, every time. I don't know why but I think it has something to do with model capacity.
Saying "Nobody knows what they are doing" is just a convenient excuse to justify terrible engineering. There is a fine line between scientific trial-and-error and proud, brute-force ignorance.Let's be clear about your "frontier":Blind Gambling: When independent labs don't understand the underlying mathematics or hardware physics, they just throw data at a wall and pray to the loss curve. That is digital alchemy, not science.The Loop: Instead of fixing structural bottlenecks or learning non-linear dynamics, people just brute-force configs. Itโs the engineering equivalent of a cat grooming itself because it has nothing else to do.Zero Legacy: This unscientific approach is why the ecosystem is flooded with overfitted, hollow checkpoints that break down outside of their strict test sets.You aren't advancing the frontier; you are just polluting the platform because you refuse to open a textbook. Brute force has hit a physical wall. True innovation requires cognitive architecture and actual engineering, not just romanticizing failure. ๐ซต๐คก
Lol. AI is being used by the very trolls that try to destroy it. Quite ironic. Quality bait though.
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
Basically, they nerfed their models, then they lower limits without any transparency. Finally, they accused their users of being using their models the wrong way. Gave their free users luna by default and nerfed back to 5-hours limit plus users.
A shared byte model makes the most sense when tokenization inefficiency is already high; for well-tokenized languages, the sequence expansion may still outweigh the benefits unless the SSM is efficient enough to absorb it. But even then, SSM being a real alternative is a huge deal compared to years ago.
1. Knowledge that deeper layers train smoothly.
2. Knowledge that Transformers work but is quadratic on sequence length.
3. Knowledge that SSMs work even better. Numerically unstable sometimes.
4. Speculative-decoding.
5. Open high-quality data.
6. Knowledge that KD works.
It slowly feels like is no longer a bad idea.
Interesting
a little bit skeptic about 1m parameters though

