context
128
ctx:discord/blah/general/128Source document
full textgeneral-128
text/plain3 KB
doc:agent/general-128/c4588bce-f1f2-4a72-896f-b209a04b555e[2026-04-11 05:00] xenonfun: yeah Nemo Cascade 2 was quite good. at 63GB at full it was usable on the 96GB max tho I was screwing around and crashed machine, at 8-bit no issues. Gemma4 26B seemed to be sweet spot, tho was still a little early on decoders, was having thought looping issues and such, been a few days probably fixed. [2026-04-11 05:01] girvo: Yeah its way more stable under latest llama.cpp in my testing now. Gemma 4 is gonna be a big slow burn i thnk [2026-04-11 05:02] xenonfun: well these new JANG quants are pretty badass MLX only seutp, he has it down to 19GB and uncensored with effectively no loss in quality. [2026-04-11 05:03] xenonfun: the GGUF stuff takes a decent hit on M-series stuff, tho the tooling to covert things over to mlx flavor is quite good. they added Gemma4 style networks in day or two across a bunch of tooling [2026-04-11 05:07] girvo: i'm trying to fine tune Gemma 4 for a work task (the E4B model I think), natural langauge to "tql" a query/search/filter langauge, its fun as lol [2026-04-11 07:52] ajaxdavis: google meets in 45 [2026-04-11 08:09] ajaxdavis: https://github.com/AlexsJones/llmfit [2026-04-11 08:22] ajaxdavis: https://meet.google.com/scv-nbnx-hxt solving world peace [2026-04-11 08:42] ajaxdavis: <@823468778704076810> p.s. you are meant to be the honorable guest so make an appearance [2026-04-11 09:51] girvo: <@806444151422976035> feck was having dinner with the fam [2026-04-11 09:51] ajaxdavis: we on and will be all night [2026-04-11 10:16] ajaxdavis: (files: sentence_woman_osman_slower_deeper.wav) [2026-04-11 10:23] ajaxdavis: https://github.com/thomasdavis/alpha2 [2026-04-11 11:02] ajaxdavis: <@1211062099137265723> [2026-04-11 11:02] ajaxdavis: (files: Screenshot_from_2026-04-11_21-01-47.png) [2026-04-11 11:07] ajaxdavis: (files: Screenshot_from_2026-04-11_21-07-19.png) [2026-04-11 11:25] ajaxdavis: https://meet.google.com/scv-nbnx-hxt working [2026-04-11 12:22] girvo: <@806444151422976035> you nerd sniped me, and now i'm looking to see if I can train an FP4 nearly-entirely 4-bit model pipeline lol > The interesting question is whether you could train a small model that's designed for FP4 from the ground up — architecture choices that play nicely with the format's logarithmic distribution, weight initializations that stay within the representable range, maybe even loss functions that penalize weights drifting into poorly-represented regions. That's the kind of thing that would produce genuine insights, because the community consensus right now is basically "train big in BF16, compress after" and very few people are questioning whether that's the only viable path. [2026-04-11 12:36] ajaxdavis: sounds like fucking fun. watching claude try resolve oom errors by jumping back and fourth before fp16/fp32/bf16 was depressing, fp4 sounds like another lay of extra fuckery, report back!
Facts in this context
Grouped by subject. Each subject links to its full article.
Ajaxdavis25 factsex:ajaxdavis
| addressesUser | User 1211062099137265723 |
| addressesUser | User 806444151422976035 |
| assessesActivityAs | Fp4 Training Pipeline |
| performedAction | Nerve Snipe Event |
| rdfs:label | ajaxdavis |
| rdf:type | User |
| requestsActionFrom | User 823468778704076810 |
| requestsReport | Girvo |
| saidAtTime | 2026-04-11 07:52 |
| saidAtTime | 2026-04-11 12:36 |
| saidAtTime | 2026-04-11 11:25 |
| saidAtTime | 2026-04-11 11:02 |
| saidAtTime | 2026-04-11 10:23 |