context
70
ctx:discord/blah/safiersemantics/70Source document
full textsafiersemantics-70
text/plain3 KB
doc:agent/safiersemantics-70/dbacde78-f635-4864-93c8-c2425e32c560[2026-02-19 20:25] xenonfun: model-ds being trained, asked it to optimize just on this training set what can be done without blowing out my 24GB limit and not exhausting the model from not enough data. (files: Screenshot_2026-02-19_at_3.23.32_PM.png) [2026-02-20 01:58] xenonfun: ``` Early stopping: val loss hasn't improved for 30 eval intervals (750 iters). Best val loss was 3.7141. Saving checkpoint and stopping. Total time: 25529.5s | Avg: 1955ms/iter ✓ Saved checkpoint.bin (iter 12300) ✓ Saved checkpoint_best.bin (best val loss 3.7141) Estimating final loss... Final train loss: 3.6974 (started 7.3126) Final val loss: 3.7805 (ppl 43.8, started 7.3126) === Generation After Training === Prompt: "ROMEO:" Come and the Elizard David was realize her parties. The Clay concluded squenance of the barticies of words from laid his head and at the hug, the centure of the most softling until it was sworked to me being the comforts, but a confliction, and even low buts on wagged Prompt: "To be or not to be" tter to hear it with me.” “A good fellow?” “Nate, for such false.” “I don’t asked my lord, but I got my heart of all.” “And then, my sister,” returned Mrs. Groarth. “How if you live no one woman do more motions to the commune-branded sympath Prompt: "Once upon a time" . That the father and half-varred by hoped single corner to the graptiles spected his character. On passed on the first-foured bows and fashioning-stairs in drill table andok in itself (the walls to the door, whom he were stood gradually close upon bidd ``` [2026-02-20 02:01] ajaxdavis: that looks like progress [2026-02-20 02:05] xenonfun: yeah the v2 cleaner data, there is a v3 of data I want to rerun. testing colab varient of training. also got hot reload on the hugging face so can update model weights on the fly for all future deployments [2026-02-20 02:06] xenonfun: the T4 on colab might be slower than the m4 just bit less memory pressure, don't like using the colab but its mostly scripts that I should be able to use on anything with card [2026-02-20 02:08] xenonfun: it will finish in another 30m here, want to compare to randy-s with the less filtered dataset. [2026-02-20 02:11] ajaxdavis: claude still exceeds my expectations often, this is gonna push it a bit harder (files: Screenshot_from_2026-02-20_12-10-58.png) [2026-02-20 04:39] lisamegawatts: nice 🙂 i am working on optimizing training data before i even touch model, [2026-02-20 04:53] xenonfun: well 500K worth of gists in Rust were already overtraining. recommends I try fine tuning the deep-small model instead will get better output, so gonna try that. (files: Screenshot_2026-02-19_at_11.50.59_PM.png) [2026-02-20 17:01] xenonfun: well finetune on rust was just randomness, tho also found out my rust trainer wasn't using dropout, now is. Running the `deep small` model again with more cleaned up data and some texts removed/added to have more consistency in text base. (files: Screenshot_2026-02-20_at_11.59.42_AM.png)
Facts in this context
Grouped by subject. Each subject links to its full article.
Log Entry 18 factsex:log-entry-1
| hasSpeaker | Xenonfun |
| hasTimestamp | 2026-02-19 20:25 |
| mentionsConstraint | Data Sufficiency Constraint |
| mentionsConstraint | Memory Limit |
| mentionsGoal | Optimization Goal |
| mentionsModel | Model Ds |
| rdf:type | Log Entry |
| referencesFile | Screenshot 1 |
Training Log Output8 factsex:training-log-output
| reportsDuration | Total Training Time |
| reportsEvent | Early Stopping Event |
| reportsFileCreation | Checkpoint Bin |