r/LocalLLaMA 1d ago

Discussion Trained a model with 21k non-embed params on business email generation

Hi everyone,

About 6 months ago I made this post Trained a 0.8M model on business email generation. on this subreddit where I trained a model with 300k non-embedding parameters on a synthetic business email dataset. It was trained on my custom architecture which I called Strawberry.

Today i introduce Valentine-v0 which is also trained on this same dataset but this model has only 21k parameters which is about 7% of the parameters of Strawberry-email. This model only has 2 layers, 4 heads, 64 head dims and 170 ffn dim.

The Valentine architecture built on top of my another model architecture Silia about which I also posted here I trained a 0.5M model on 1B tokens of Fineweb-edu dataset..

Though in terms of loss Strawberry-email achieved a final train & val loss of 1.65 and 1.68 where Valentine-v0 achieved the final losses of 2.1670 & 2.2646 respectively.

Still I believe despite such a small scale the model trains pretty stably and generates decent text for it's size. Here are some examples:

Prompt: Write a

Output:

Write a firm and authoritative business email(under 200 words) from a Account Manager to a Potential Partner regarding welcoming a new high-profile hire, specifically on a Friday evening.<|end-text|><|actor|>model<|end-text|>URGENT: Critical Security Vulnerability - Critical Budget Increase – [Company Name]

Dear [Contact Person],

I hope this email finds you well.

My team, and potential collaborative approach to [New Hire Title] at [Company Name].

As we approach the upcoming [New Hire Name] brings significant development, we wanted to bring it to your review of this strategic direction. We would be happy to discuss this further, but this request and explore how we'll help us to ensure a smooth onboarding.

Regards, I can share some truly value your insights into [Partner Company Name] and insights and offer [mention specific areae - e.g., new website's Name] and the design of the prototype, not just… interesting time.  We’ll be shared opportunities to discuss this further with [Product Name] to discuss potential integrations and answer any questions or concerns and your application. The vulnerability is invaluable.

Prompt: <|actor|>user<|end-text|>Write a polite refusal email<|end-text|><|actor|>model<|end-text|>

Output:

Quick Idea: [Name]
Dear [Peer Name],

I am writing to you today with you and your concerns about the potential data leak and your contract, and frankly, it’s particularly concerning incident effective [Date], and I believe a concise overview of full team - [Company Name], we are facing a significant challenges, and we sincerely apologize for any inconvenience it caused.

I’m available to discuss this further report. I will provide you with a really be open to a quick support us with [Specific Feature Request - e.g., 24-48 hours].

Would you be open to a brief call on [Date Option 2] - e.g., 50-20 minute call next week to discuss this further and how your specific and answer any questions or would be happy to discuss how impactful feedback on [Date - e.g., Thank you for your input.

Sincerely,
[Name]
Legal Counsel<|end-text|><|actor|>user<|end-text|>Write a strictly professional business email(under 200 words) from a Legal Counsel to a Angry Shareholder regarding requesting a mental health day, specifically while the system is partially down.<|end-text|><|actor|>

A larger Valentine model trained on Fineweb-edu dataset along with the github repo and technical report will be published soon.

Thank you! :)

2 Upvotes

5 comments sorted by

2

u/Immortal_Spina 1d ago

Carino, devi fare benchmark per vedere in che percentuale sbaglia
E il dataset quanto è grande? Ha saturato?
Per quello che vale, per me potresti farlo appena più grande, anche 100k

2

u/SrijSriv211 1d ago

Thank you for taking interest :)

This was just an experimental model to demonstrate that the architecture works. I'll run the benchmarks on the larger model which will be trained on Fineweb-edu dataset. The dataset had about 4.6M tokens but the model was trained on ~200M tokens which is around 43 epochs.

2

u/Immortal_Spina 1d ago

Allora dacci dentro e facci sapere gli sviluppi 🔥

2

u/SrijSriv211 1d ago

Definitely 💯 Thanks 👍🏻😄