boccaccioAI is a 100% Italian LLM with 700M parameters, built entirely from scratch by De Lauretis Tech. No pre-trained weights, no external wrappers, just a custom tokenizer, architecture, and weights trained on Italian text. Think of it as an elementary school child: it might hallucinate details, but its grammar, syntax, and flow are flawlessly Italian. A proof-of-concept proving that sovereign LLMs are possible at scale. Running on on-demand serverless GPUs. Try it now, using italian :D!
No reviews yetBe the first to leave a review for boccaccioAI - 100% Italian LLM
Maker
📌
Hi Product Hunt!
I’m Lorenzo, founder of De Lauretis Tech, and today we’re incredibly excited to share boccaccioAI with you.
Most LLMs today are fine-tuned versions of English-centric models. We wanted to see if we could build a truly independent, sovereign Italian LLM completely from scratch, and we did.
What makes boccaccioAI unique:
- 100% Custom: Custom tokenizer, custom architecture, and custom weights trained from a random initialization purely on Italian data. No third-party pre-training.
- 700M Parameters: It's small, lightweight, and runs on serverless GPUs (which spin up on-demand; expect a ~20s cold start on your first prompt!).
- The "Elementary School" Phase: Since it's small, it will hallucinate facts (it's still learning about the world!), but pay attention to its flawless Italian grammar, natural syntax, and flow. The foundation is solid.
What’s next?
This 700M model is our proof of concept. We are already preparing the training pipeline for a 7B parameter model to bring true factual knowledge, reasoning, and depth to Italian AI.
We’d love to hear your thoughts, feedback, and questions. Give it a spin and let us know what you think of the generation quality! Obviously, you will have to write in Italian 😄
Report
love seeing a sovereign italian model out there, the language quality sounds really promising. one thing that would help me as a user: a side-by-side compare button where i can paste the same prompt into boccaccioai and chatgpt to see how the italian phrasing differs. would make the "native fluency" angle way more tangible for non-technical folks like me checking it out.
Report
Maker
@oztrurk73771 This is a fantastic idea! A side-by-side comparison focusing purely on the phrasing and flow (while obviously ignoring the factual knowledge gap between our 700M model and a massive trillion-parameter giant) would be a brilliant way to showcase the native vs 'translated' feel. I'm adding this to our feature backlog right now. Thanks for checking it out!
Report
Tried asking it about a regional Tuscan recipe and the response was fully in proper Italian, even down to the little flourishes of phrasing. Pretty cool proof of concept, though it did make up an ingredient or two.
Report
Maker
@sabasyaren82598 Thank you! I love that you tested it on regional recipes. It might have just invented a new avant-garde Tuscan fusion dish! But hearing that the nuances and flourishes of the language came through naturally is exactly the feedback we were hoping for. The actual culinary knowledge will definitely improve when we scale up!
Report
Tried asking it about a Tuscan recipe and the response genuinely sounded like something a nonna would say, grammar and all. Wild that this came out of a 700M model built from scratch.
Report
Maker
@sedanur1673693 The 'nonna' comparison is the ultimate compliment! Capturing that authentic, conversational tone is exactly why we decided to build the tokenizer and model entirely from zero instead of fine-tuning an English base. Thank you for putting it to the test.
Report
Gave it a quick spin in Italian and the phrasing genuinely feels native, not that translated vibe most small models have. Love seeing a sovereign build like this, even at 700M params the grammar really does hold up.
Report
Maker
@erdihbhf Spot on. Eliminating that exact 'translated' feeling is the primary reason we built the tokenizer and architecture natively rather than relying on existing bases. It is highly rewarding to hear that the 700M foundation holds up so well in your tests.
Report
Maker
Hi Product Hunt community,
Thank you to everyone who is testing boccaccioAI, leaving comments, and supporting our launch right now.
The live feedback we are receiving is exactly what we are aiming for. As many of you are noticing, while a 700M parameter model inevitably hallucinates facts, the grammar, syntax, and phrasing are flawlessly native. This validates our core hypothesis: building an LLM completely from the ground up, with a custom native tokenizer and architecture, is the right path to avoid the "translated" feel of mainstream models.
Seeing this custom foundation hold up in real-time is our primary goal today. Your ongoing tests and active discussions are giving us great momentum as we prepare for our next major milestone: scaling the architecture and training pipeline to a 7 Billion parameter model. That is where we will bridge the gap between knowing how to speak and actually knowing about the world.
If you believe in the future of European sovereign AI, or want to explore technical synergies or investment opportunities with De Lauretis Tech, I would love to connect.
Please keep testing the model, exploring its limits, and sharing your feedback on the linguistic flow. Thank you for the continuous support!
Report
Built entirely from scratch is genuinely impressive, love seeing this kind of sovereign AI work. Threw a few Italian prompts at it and the phrasing feels natural even when it fumbles facts.
Report
Maker
@pustundal56720 Thank you! Proving that a ground-up approach results in better linguistic flow was our main objective today. We are perfectly fine with it fumbling facts at 700M parameters, as long as it does so in flawless Italian.
love seeing a sovereign italian model out there, the language quality sounds really promising. one thing that would help me as a user: a side-by-side compare button where i can paste the same prompt into boccaccioai and chatgpt to see how the italian phrasing differs. would make the "native fluency" angle way more tangible for non-technical folks like me checking it out.
@oztrurk73771 This is a fantastic idea! A side-by-side comparison focusing purely on the phrasing and flow (while obviously ignoring the factual knowledge gap between our 700M model and a massive trillion-parameter giant) would be a brilliant way to showcase the native vs 'translated' feel. I'm adding this to our feature backlog right now. Thanks for checking it out!
Tried asking it about a regional Tuscan recipe and the response was fully in proper Italian, even down to the little flourishes of phrasing. Pretty cool proof of concept, though it did make up an ingredient or two.
@sabasyaren82598 Thank you! I love that you tested it on regional recipes. It might have just invented a new avant-garde Tuscan fusion dish! But hearing that the nuances and flourishes of the language came through naturally is exactly the feedback we were hoping for. The actual culinary knowledge will definitely improve when we scale up!
Tried asking it about a Tuscan recipe and the response genuinely sounded like something a nonna would say, grammar and all. Wild that this came out of a 700M model built from scratch.
@sedanur1673693 The 'nonna' comparison is the ultimate compliment! Capturing that authentic, conversational tone is exactly why we decided to build the tokenizer and model entirely from zero instead of fine-tuning an English base. Thank you for putting it to the test.
Gave it a quick spin in Italian and the phrasing genuinely feels native, not that translated vibe most small models have. Love seeing a sovereign build like this, even at 700M params the grammar really does hold up.
@erdihbhf Spot on. Eliminating that exact 'translated' feeling is the primary reason we built the tokenizer and architecture natively rather than relying on existing bases. It is highly rewarding to hear that the 700M foundation holds up so well in your tests.
Hi Product Hunt community,
Thank you to everyone who is testing boccaccioAI, leaving comments, and supporting our launch right now.
The live feedback we are receiving is exactly what we are aiming for. As many of you are noticing, while a 700M parameter model inevitably hallucinates facts, the grammar, syntax, and phrasing are flawlessly native. This validates our core hypothesis: building an LLM completely from the ground up, with a custom native tokenizer and architecture, is the right path to avoid the "translated" feel of mainstream models.
Seeing this custom foundation hold up in real-time is our primary goal today. Your ongoing tests and active discussions are giving us great momentum as we prepare for our next major milestone: scaling the architecture and training pipeline to a 7 Billion parameter model. That is where we will bridge the gap between knowing how to speak and actually knowing about the world.
If you believe in the future of European sovereign AI, or want to explore technical synergies or investment opportunities with De Lauretis Tech, I would love to connect.
Please keep testing the model, exploring its limits, and sharing your feedback on the linguistic flow. Thank you for the continuous support!
Built entirely from scratch is genuinely impressive, love seeing this kind of sovereign AI work. Threw a few Italian prompts at it and the phrasing feels natural even when it fumbles facts.
@pustundal56720 Thank you! Proving that a ground-up approach results in better linguistic flow was our main objective today. We are perfectly fine with it fumbling facts at 700M parameters, as long as it does so in flawless Italian.