Papers
arxiv:2603.03543

Tucano 2 Cool: Better Open Source LLMs for Portuguese

Published on Mar 3
Authors:
,
,

Abstract

Tucano 2 is an open-source suite of Portuguese language models with varied parameter counts, enhanced datasets, and comprehensive evaluation methods for improved language understanding and generation.

We present Tucano 2, a fully open suite of large language models (LLMs) with 0.5-3.7 billion parameters, designed to address certain gaps in open-source development for Portuguese LLMs. Following our previous works, we now extend our dataset, GigaVerbo-v2, to a new degree of quality and scale, while also introducing a new synthetic dataset, GigaVerbo-v2 Synth, aimed at filling missing gaps in GigaVerbo-v2, and two post-training datasets, GigaVerbo-v2 SFT and GigaVerbo-v2 Preferences, that allow Portuguese LLMs to be trained in domains like retrieval augmented generation, coding, tool use, chain-of-thought reasoning, and many other domains of interest. Through extensive ablation studies, we design both pretraining and continual pretraining recipes for the Tucano 2 suite (Base, Instruct, and Think), which achieve state-of-the-art performance on several Portuguese-language modeling benchmarks. We also extend and refine the evaluation harness introduced in our earlier work, yielding a comprehensive evaluation suite that provides strong signals across different pretraining, continual pretraining, and post-training regimes. All artifacts associated with Tucano 2 are openly released, including training recipes, logs, and source code, ensuring that our work is reproducible, accessible, and extendable by the broader Portuguese NLP community.

Community

•
This comment has been hidden (marked as Abuse)
Paper author
•
This comment has been hidden (marked as Off-Topic)
•
This comment has been hidden (marked as Abuse)

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2603.03543
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 26

Browse 26 models citing this paper

Datasets citing this paper 10

Browse 10 datasets citing this paper

Spaces citing this paper 4

Collections including this paper 2