First time comment user here

#6
by Ashwin097 - opened

works well on 2070s maxq, lags and struggle after reaching 10-13k context, cant complain since it took many chats to reach there , using like whatsapp chat in low end , this maybe perfect , logged in just to comment , first comment for the models to this , whats this slow burn mentioned in other LLM , does this also have by the way

Glad to hear!

Also, for slow-burn specifically, Angelic_Eclipse_12B might be even better :)

I'd found MOE models tend to be quite a bit faster (26B/35B); That's probably the only reason i haven't used 12B/27B/31B models much especially lately.

Though when you're running on mostly CPU you can go to like 70B and get nearly the same speed (like 1t/s). Some good 70B's out there. (1-2t/s is on par with a human, as in get a RP response in about 10 minutes. Slower then 1t/s i consider unusable).

Some good 70B's out there.

Got any personal 70B recs for 2026? Ones with decent smarts, for use as a GM engine or similar.
Sorry to derail.

Some good 70B's out there.

Got any personal 70B recs for 2026? Ones with decent smarts, for use as a GM engine or similar.
Sorry to derail.

you gotta mention your hardware specs,

Ashwin097 changed discussion status to closed
Ashwin097 changed discussion status to open
Ashwin097 changed discussion status to closed

Some good 70B's out there.

Got any personal 70B recs for 2026? Ones with decent smarts, for use as a GM engine or similar.

There will be an Impish_LLAMA_70B one day, but that day is not today πŸ˜‰

There will be an Impish_LLAMA_70B one day, but that day is not today πŸ˜‰

Sounds promising.

Got any personal 70B recs for 2026? Ones with decent smarts, for use as a GM engine or similar.

I got a list of recommended models i add to from time to time - https://huggingface.co/collections/yano2mch/potential-models-story-rp-technical

But if I'd have to recommend a singular 70B model, it would be Darkhn-Quants-2's Animus V12.5. Even heavily quantized it performed very well (at least from memory).

There will be an Impish_LLAMA_70B one day, but that day is not today πŸ˜‰

Sounds promising.

Got any personal 70B recs for 2026? Ones with decent smarts, for use as a GM engine or similar.

I got a list of recommended models i add to from time to time - https://huggingface.co/collections/yano2mch/potential-models-story-rp-technical

But if I'd have to recommend a singular 70B model, it would be Darkhn-Quants-2's Animus V12.5. Even heavily quantized it performed very well.

am i the only here with potato laptop ? , you all casually say 30 50 70B meanwhile i struggle to try a 12 B model itself 😭

Ashwin097 changed discussion status to open

am i the only here with a potato laptop? , you all casually say 30 50 70B meanwhile i struggle to try a 12B model itself 😭

Most laptops are pretty weak unless they are gaming laptops. I updated my system to 128Gb RAM (from 32Gb) JUST before the ram prices went up. As long as you have enough RAM you can run almost any model (though speed via CPU is another matter entirely, though MOE models will be faster). Think at Q2 i had a 70B model down to ~28Gb, and some perform rather well even at Q2.

But run what you can i guess; Least till RAM & GPUs are cheaper.

edit: Actually, you might consider the bonsai models. 1 and 2bit 27b. Comes to a modest 5Gb and 9Gb. Yeah probably not as RP tuned, but may work for you.

am i the only here with potato laptop ? , you all casually say 30 50 70B meanwhile i struggle to try a 12 B model itself 😭

If you're curious to try out 70Bs, then MegaNovaAI serves Sapphira, Nevoria and Euryale on the free tier.
Out of the 1-2 shot testing I've done, Sapphira was the one managed to give me a chuckle, if you're looking for one that can play rough, but afaik Sicarius' Pepe ought to outdo it on those metrics. My own worry is that the Pepe personality might make it a poor neutral narrator/GM, but I haven't actually tested it for that.
2026-06-02_10-18-36

Not Bloodmoon, but last I've tested Impish Nemo a while back, it's been able to use refreshingly restrained and clear narration, if a tad repetitive, compared to the slog of purple purple from various Mistral Small 24B tune competitors. But a heavy sys prompt might've been at fault.
Also, heads above the competition when asked for a 'completely incomprehensible accent'!
2026-02-22_18-30-22

if a tad repetitive...

I remember trying out a model i am sure it was a 8b or 12b model, after the 7th 'bursts into a fractal pixel galaxy before reforming in finer detail' or the like i was done (Mind you it was an AI hacking my phone and holding it hostage, hoping to become a real girl or something...)

Too much repetition and low quality are the fastest ways to lose interest in a model or RP. Or fitting into a niche scenario it has no idea how to get out of or progress so it just restates the current state of things with no progress repeatedly.

Sign up or log in to comment