The Corpus as a Public Good
By Anushka Appala and Dr. Janio Rosales
When a linguist at the Academia de Lenguas Mayas checks how NaciluzIA renders a benefits question into Q'eqchi' — correcting a verb here, flagging a regional spelling there, noting that a phrase reads as bureaucratic when it should read as plain — something quietly durable is being made. Not just a better answer for the next citizen who calls, though that matters. Each correction is deposited into a growing, validated record of how a Mayan language actually works in the setting of public service. Over time, those deposits become an asset that outlives any single program: a verified Mayan-language corpus for the State.
We think this is one of the most important things NaciluzIA will produce, and it is worth explaining why it cannot be shortcut.
Why Spanish to Q'eqchi' is not Spanish to French
It is tempting to assume machine translation is a solved problem — point a model at a language pair and let it run. That assumption holds only where the world has already done the model's homework. Spanish to French rests on centuries of parallel text: treaties, novels, newspapers, legal codes, dictionaries, standardized spelling, millions of aligned sentence pairs. The data is deep, clean, and abundant.
Spanish to Q'eqchi' is a different problem in kind, not degree. The digitized parallel text is sparse. Spelling is not fully standardized, so the same word appears several defensible ways. Much of the language lives in oral tradition rather than written record, which is exactly where large text-trained models are weakest. And "Q'eqchi'" is not monolithic — usage shifts by region, so a phrasing that is natural in one municipality can land wrong in the next. A model trained on what little exists online will be confidently, invisibly mistaken in ways no automated metric reliably catches.
This is not a reason to avoid the work. It is the reason the work must be done responsibly — with people who speak the language in the loop.
The human-validation loop, and what it leaves behind
So NaciluzIA treats indigenous-language understanding as a partnership, not an API call. Every Mayan-language interaction is validated with the Academia de Lenguas Mayas: native speakers and linguists review what the system understood and how it responded, and their judgments feed back into it. The AI proposes; the people who own the language decide whether it is right. This is the same principle that governs the rest of the platform — AI recommends, humans decide — applied to meaning itself.
The loop does two things at once. In the moment, it protects the citizen: no one receives a garbled or subtly disrespectful rendering of a right they are trying to claim. And over time, it accumulates. Every validated exchange is a small, checked contribution to a corpus of real public-service language — questions people actually ask, answers that actually landed, terminology confirmed by the institution charged with stewarding these languages. That corpus does not belong to a vendor. Like the rest of NaciluzIA, it is meant to be an open, public asset — built with public participation, held for the public interest.
An asset that compounds, at national scale
Consider the reach this loop can eventually serve. The four largest Mayan languages together account for roughly 2.5 million people — Q'eqchi' about 1.3 million, K'iche' about 1.1 million, Mam roughly 600,000, and Kaqchikel about 410,000 native speakers (2019 census / Ethnologue). A validated corpus for even these four is not a research curiosity; it is infrastructure for a large share of a nation, and a foundation the country does not currently possess.
Once it exists, it compounds. A corpus validated for pensions helps the next program reach the same speakers faster. What was built for one ministry can serve health, education, land records, courts. And because it is documented and open rather than locked inside a proprietary system, it becomes a resource for teachers, researchers, and other public institutions — a standing contribution to the languages themselves, not a byproduct discarded when a contract ends.
There is a lesson here about the right pace for this kind of work. The responsible path — humans in the loop, validation by the institution that owns the language — is slower than scraping the web and shipping. But it is the only path that both protects the citizen today and leaves something valuable behind tomorrow. We are not just translating a benefits form. We are helping build, carefully and in the open, a public record of how Guatemala's languages carry the business of the State — so that dignity in one's own language is not a promise renewed program by program, but an asset the country keeps.
Cada decisión, a la luz.
— Anushka Appala and Dr. Janio Rosales