Is AI Diluting Patois and Creole?

The Culture Machine · The Language Question

Sixteen territories, sixteen languages spoken at home, and only two of those languages carrying official status anywhere. Artificial intelligence has arrived in the middle of that arrangement and is doing two opposite things at once. It is producing a blended, English-leaning imitation of Caribbean creoles, and it is building the first permanent records those languages have ever had. Sixteen territories, two answers, and one decision the region has not yet made.

By the Caribbean AI Newsletter 18 August 2026 16 min read

In Garifuna, men and women use different words for the same thing. The men's set draws mostly from Carib, the women's from Arawak, and the split has survived four centuries, a forced deportation from Saint Vincent, and resettlement across Belize, Honduras, Guatemala and Nicaragua. UNESCO recognizes Garifuna as a Masterpiece of the Oral and Intangible Heritage of Humanity and lists it as vulnerable in its Atlas of the World's Languages in Danger.

Now ask a language model to write Garifuna. It has no signal telling it which words belong to whom. It will average them.

That averaging, repeated across every Caribbean language, is what people mean when they ask whether AI is diluting Patois and Creole. It is a fair question and it has a real answer, which is not the one either side of the argument usually wants.

AI is doing both at the same time. It is thinning Caribbean creoles by generating a blended, English-leaning version of them, and it is preserving them by creating the first machine-readable records many of these languages have ever had. Which effect wins depends on whether Caribbean institutions build the training data or leave foreign models to guess.

Machines guess when the record runs thin

This is not a story with a villain. The same mechanism drives both outcomes: machines learn a language from whatever record of it exists. Where the record is thin, they invent. Where it is good, they preserve.

At a glance
The case that AI is thinning our languages, and the case that it is rescuing them
Both columns are true right now, in different territories, for different languages.

The thinning caseThe machine averages what it does not know

  • It writes creole it cannot speakGPT-4 reads Guyanese Creolese at 29.80 BLEU and writes it at 1.35. What it produces is English wearing an accent.
  • The big creole absorbs the small onesJamaican Patois has the most data, so models reach for it when they meet Bajan, Vincentian or Antiguan speech.
  • It flattens the continuumCaribbean speech runs on a scale from broad creole to standard English. Models collapse the scale and lose code-switching entirely.
  • It teaches you to correct yourselfSpeakers change how they talk so machines will understand, and the machines then learn from the corrected version.

The rescue caseOral languages are getting written down

  • First records everKreyòl-MT produced the first parallel text that has ever existed for 21 creole languages, out of 41 it covers.
  • Three of ours are now listedHaitian Creole, Jamaican Patois and Papiamento are all in Google Translate.
  • Small data worksA 400 million parameter model fine-tuned on fewer than 2,000 sentences writes Creolese nine times better than GPT-4.
  • Entry is getting cheaperMeta's 2025 speech system can be extended to a new language with a handful of recorded samples rather than a research budget.

The thinning case: what the machine gives back is not what you gave it

In 2024 a team from the University of Michigan and the University of Guyana ran the first serious translation benchmark on Creolese, the English-lexicon creole most Guyanese speak at home. They tested six systems in both directions and published the results at NAACL 2024.

GPT-4, given no examples, translated Creolese into English at 29.80 BLEU. Asked to translate English into Creolese, the same model scored 1.35.

BLEU runs from 0 to 100 and competent human translation sits in the 50s. A 29.80 is rough but workable. A 1.35 is noise with correct punctuation. Same model, same language pair, and a twenty-two-fold difference depending on which way the arrow points.

Loading the prompt with worked examples barely helped. Few-shot prompting lifted the writing score from 1.35 to 1.64 while the reading score stayed near 30. Whatever the model had picked up about Creolese, it had only learned how to decode it.

A name for the mechanism
Caliban's Autocorrect

Shakespeare gave Caliban one famous line about language: he was taught Prospero's tongue, and what he got from it was the ability to curse. Caribbean writers from George Lamming onward made him the figure of the colonized speaker, holding a language that was never built for him.

The machine version is quieter and more efficient. You write in Patois. The system understands you well enough, and answers in English. Next time, you write in English first, because that is faster. The system logs the English version as the data, learns from it, and the next speaker gets a slightly worse creole model than you did.

No single exchange does damage. The direction of travel is the point.

The big creole absorbs the small ones

Jamaican Patois has more digital text than every other English-lexicon Caribbean creole combined, because of reggae, dancehall, film, and a diaspora that writes online. That makes it the shape a model reaches for whenever it meets Caribbean speech it does not recognize.

A Bajan sentence, a Vincentian sentence and an Antiguan sentence are all closer to Jamaican than to standard English, so the model pulls them toward Jamaican. The output sounds Caribbean. It is not Barbadian, and a Barbadian will tell you so in about four seconds.

This pattern has a documented precedent that has nothing to do with AI. In Tropical Tongues, published by the University of North Carolina Press in 2018, Jennifer Carolina Gómez Menjívar and William Salmon traced what happened in coastal Belize after independence in 1981. Kriol rose to the status of a national language. Mopan and Garifuna, spoken in the same communities, attenuated. A dominant creole can crowd out a smaller neighbour without anyone intending it.

Machine learning does that arithmetic faster and at national scale, because the model is explicitly optimizing for whatever it has most of.

Code-switching disappears

Caribbean speech is not two languages with a wall between them. It runs on a scale, from broad creole through the middle registers to standard English, and speakers move along it constantly depending on who they are talking to. A Trinidadian at a funeral, in a meeting and on a WhatsApp group is using three different points on that scale within an hour.

Models have almost no training data for the middle of that scale, and none at all for the switch itself. What they produce is a single fixed register near the broad end of the scale, because that is the version written down for effect in song lyrics and comedy. A model asked for Trinidadian will give you carnival, not a meeting.

Choosing one standard turns the other varieties into spelling errors

Fixing any of this means choosing one written standard per language. Cassidy-JLU for Jamaican. Cave-GLU for Creolese. The National Kriol Council's system for Belizean Kriol, which already has a published dictionary and a grammar. That choice is technically necessary and it is not neutral. Whichever variety the standard was built from becomes the machine-readable version, and the others become spelling errors.

Creoles have always changed by contact

Creole languages were born from contact. Jamaican Patois exists because West African languages met English on plantations, and it has absorbed Spanish, Taino, Rastafarian coinages, American slang and Nigerian internet English since. Treating every change as dilution would mean freezing a living language at whichever year the observer happens to prefer.

This Newsletter does not have a clean answer to that. The most defensible line is narrow: change driven by speakers is language, change driven by a system averaging over an absence of data is not, and the second one is what the 1.35 measures.

The rescue case: oral languages are being written down for the first time

The other half of this is real, and Caribbean commentary tends to skip it.

In 2024 a team from Johns Hopkins and sixteen co-authors published Kreyòl-MT, covering 41 creole languages across 172 translation directions with 14.5 million parallel sentences, 11.6 million of which were released publicly. For 21 of those 41 languages, it produced the first parallel text that has ever existed. Half the creole languages in that set had no machine-readable record at all until two years ago.

In June 2024 Google added 110 languages to Translate in a single release, taking its total to 243. Jamaican Patois and Papiamento were both on that list. Haitian Creole had been there for years. Three Caribbean creoles are now served by the most-used translation tool on earth, and for a Haitian in a Miami emergency room or a Kingston vendor with a foreign customer, a rough translation beats none.

Meta's Omnilingual ASR, released in November 2025, covers more than 1,600 languages for speech, over 500 of which no system had ever transcribed. The design decision that matters for small territories is that communities can add an unserved language with a handful of their own recorded samples rather than a research grant.

And in Guyana, a Creolese-speaking assistant called IRIS answers citizens over WhatsApp, built on a corpus of 2,373 sentences.

A name for the opportunity
The Anansi Archive

Anansi crossed the Atlantic in memory. No ship's manifest carried him, no book preserved him, and he survived because each generation retold him to the next. The Caribbean's languages have travelled the same way: spoken, remembered, and never written down at scale.

Transcription changes the medium of survival. A recorded, transcribed corpus is the first form in which these languages outlive the people who speak them. That is a genuinely new thing in four hundred years, and it is available now to any territory willing to record ten hours and transcribe them properly.

There is a hard version of this argument too. Roughly 40 percent of the world's 7,000 languages are endangered, on the figure Google cited when it relaunched its Woolaroo language app in March 2026. For a language with 120,000 speakers and a falling transmission rate, a machine-readable corpus decides whether the language can be relearned at all.

Exhibit 1 · Reading versus writing Guyanese Creolese
Every system tested reads Creolese better than it writes it. Fine-tuning almost closes the gap
BLEU measures translation quality from 0 to 100, where competent human translation sits in the 50s. Higher is better. The two GPT-4 bars are general-purpose prompting. The four on the right are smaller models fine-tuned on the GuyLingo corpus of fewer than 2,000 sentence pairs.
322416 80 BLEU SCORE Creole into English (reading) English into Creole (writing) 29.830.219.7 17.714.26.1 1.351.649.74 12.110.22.67 GPT-4GPT-4T5-Large BART-LargeBART-BasePegasus zero-shotfew-shotfine-tuned fine-tunedfine-tunedfine-tuned Thinning: 18x to 22x gap Rescue: under 2.3x, after fine-tuning on fewer than 2,000 sentence pairs

Source: Clarke, Daynauth, Wilkinson, Devonish and Mars, "GuyLingo: The Republic of Guyana Creole Corpora", NAACL 2024. Ratios calculated by the Caribbean AI Newsletter from the paper's Tables 3 and 4.

BART-Large, roughly 400 million parameters, fine-tuned on fewer than 2,000 Creolese sentence pairs, writes Creolese at 12.11 BLEU. GPT-4 manages 1.35. The small local model beats the large foreign one by a factor of nine, and narrows the direction gap from twenty-two-fold to under one and a half.

Nobody in this region needs to train a foundation model to make creole work. The route runs through recording, transcription and fine-tuning, and it has been walked already.

Nine of sixteen territories have nothing built for them

The regional picture is uneven in a way that follows population and diaspora size rather than need. Three Caribbean creoles are in Google Translate. Three have a published research dataset. Nine of the sixteen territories this Newsletter covers have neither.

Exhibit 2 · The regional language map
Sixteen territories, what they speak at home, and whether any AI system has been built for it
Status shows what this Newsletter could verify as of August 2026. Served means the language is supported in Google Translate. Researched means a named academic dataset or framework exists. Unmapped means neither could be found.
TerritoryWhat people speak at homeStatusWhat exists
Haiti Haitian Creole (Kreyòl Ayisyen), official alongside French, around 13 million speakers, regulated by the Akademi Kreyòl Ayisyen Served Longest-standing AI coverage of any Caribbean creole. In Google Translate and Meta's NLLB models
Jamaica Jamaican Patois (Patwa), Cassidy-JLU orthography adopted by the Jamaican Language Unit in 2002 Served Researched Added to Google Translate June 2024. JamPatoisNLI reasoning dataset, EMNLP 2022
Aruba Papiamento, official, regulated by the Papiamento Academy Foundation Served Added to Google Translate June 2024
Curaçao Papiamentu, official, shared with Bonaire Served Added to Google Translate June 2024
Guyana Creolese, mother tongue of the majority. Also Wapichan, Makushi, Wai Wai, Akawaio, Arekuna, Patamuna, Kalina, Warrau and Lokono Researched GuyLingo corpus, NAACL 2024. IRIS assistant on WhatsApp. Indigenous languages uncovered
Trinidad and Tobago Trinidadian and Tobagonian English Creole. French Creole (Patois) survives in pockets Researched Translation framework published in the Caribbean Educational Research Journal, 2026
Dominican Republic Spanish. Haitian Creole among Haitian communities. Samaná English among the descendants of nineteenth-century settlers Served Partly Spanish is fully served. The creole and Samaná English layers are not
Belize Belizean Kriol, spoken by 44.6 percent of the population at the 2010 census. Also Garifuna, Mopan, Q'eqchi', Spanish and Plautdietsch Unmapped Best writing infrastructure in the region: the National Kriol Council standardized the orthography and published a dictionary and grammar. No AI dataset built on it
Barbados Bajan (Barbadian Creole), the everyday language of most Barbadians Unmapped No Translate listing and no dedicated dataset identified
The Bahamas Bahamian Creole. Haitian Creole among Haitian-Bahamian communities Unmapped No Translate listing and no dedicated dataset identified
Saint Lucia Kwéyòl, French-lexicon, spoken by a significant majority. Jounen Kwéyòl has been held annually since 1984 Unmapped A 24-letter Kwéyòl writing system exists, shared with Dominica. No dataset identified
Dominica Kwéyòl, mutually intelligible with Saint Lucian, Martinican and Guadeloupean varieties. Also Kalinago Unmapped Antillean Creole alphabet developed in the late 1970s. No dataset identified
Grenada Grenadian Creole English. Grenadian Creole French, now spoken by few Unmapped No Translate listing and no dedicated dataset identified
Saint Vincent and the Grenadines Vincentian Creole. The Garifuna were deported from here in the 1790s, taking the language to Central America Unmapped No Translate listing and no dedicated dataset identified
Antigua and Barbuda Antiguan and Barbudan Creole, part of the Leeward Caribbean Creole group shared with Saint Kitts and Nevis, Anguilla and Montserrat Unmapped No Translate listing and no dedicated dataset identified
Suriname Sranan Tongo, around 520,000 first-language and 150,000 second-language speakers. Also Sarnami Hindustani, Javanese, Saramaccan, Ndyuka and Indigenous languages, with Dutch official Unmapped Appears in multi-creole research collections. No Suriname-led dataset identified

Not listed above, and in the same position or worse: Saint Kitts and Nevis, Montserrat, Anguilla, the British and United States Virgin Islands, the Cayman Islands, the Turks and Caicos Islands, Sint Maarten, Bonaire, Martinique, Guadeloupe and Puerto Rico. The Caribbean diaspora in New York, Toronto, London, Miami and Amsterdam speaks most of these languages too, and is where a large share of the written creole on the internet is produced.

Read the table by pattern rather than by row. Every territory in the served column has either a very large speaker population, official status for its creole, or both. Every territory in the unmapped column is small, and its language is unofficial.

Belize is the sharpest case. It has the best written infrastructure of any English-lexicon creole in the region, with a standardized orthography, a published dictionary and a grammar produced by the National Kriol Council. None of it has been turned into a training set. The work of the last thirty years is sitting one step away from being useful to a machine, and nobody has taken that step.

Four forces kept creole out of the written record

People speak creole and write English. The writing systems exist but are not taught in schools. Speakers have learned to change how they talk when a machine is listening. And creole has never been the language of courts, ministries or banks, so no institution ever had a reason to record it.

People speak creole and write English

A Jamaican speaks Patois to a taxi driver and writes English to a bank. A Vincentian argues in creole and files a report in standard English. The GuyLingo authors described the mechanism for Guyana precisely: Creolese is the mother tongue of the majority, English has traditionally been the only language children are taught to read and write, and written resources in Creolese are correspondingly scarce.

The web crawls behind every major model found almost no creole because almost no creole is typed.

The orthographies exist and sit unused

Frederic Cassidy proposed a phonemic writing system for Jamaican in 1961. The Jamaican Language Unit, set up at UWI Mona in 2002, revised it into the Cassidy-JLU system. Guyana has Cave-GLU, from George Cave's 1970 work. Belize has the National Kriol Council standard. Saint Lucia and Dominica share a 24-letter Kwéyòl system.

None of them is taught at scale. The GuyLingo team reported what that costs: their training data mixed spelling conventions, one sound appeared as several letter combinations, and translation into Creolese suffered for it. Inconsistent spelling fragments the signal.

Speakers were trained to hide how they speak

On 24 June 2025, Howard University and Google Research released the Howard African American English Dataset 1.0 through Project Elevate Black Voices: over 600 hours of African American English from speakers across 32 US states.

The finding that matters here was about the speakers rather than the model. The researchers reported that natural African American English is scarce in existing speech data because Black users have been implicitly conditioned to change their voices when talking to voice technology.

Every Jamaican who has flattened an accent for Siri produced that same absence. So did every Trinidadian who retyped a query in careful standard English after the creole version failed. Part of what is missing from the training data is the record of people hiding how they speak in order to be understood, and the systems then learn from the hidden version.

Creole was never an institutional language

The Library of Congress research guide on Caribbean creoles notes that using European languages to the exclusion of creoles in schools, courts and government has kept large parts of the population from full participation in civil society, an arrangement some scholars describe as linguistic apartheid. That predates AI by centuries.

What AI changes is the enforcement. A ministry chatbot that only works in standard English does not create the exclusion. It automates it, runs it around the clock, and attaches a service-level agreement to it. A hospital triage line, a customs declaration portal, a bank's fraud call, a hurricane alert system: each of these works in the register the region writes, not the one it speaks, and nobody involved chose that on purpose.

Dr Rossana Herrero-Martín, who coordinates the UWI Cave Hill Translation Bureau, told the UN Information Centre for the Caribbean this year that the region's under-resourced languages lack the data needed to train translation models, which reduces their value as vehicles of knowledge and culture for their own communities. She argued that building those corpora with language communities rather than about them should be treated as an act of cultural justice, and used the phrase community language sovereignty for what has to be protected in the process.

Small local models already beat GPT-4 at writing creole

Guyana · 2024
A national assistant built on 2,373 sentences

Christopher Clarke, Roland Daynauth, Charlene Wilkinson, Hubert Devonish and Jason Mars built GuyLingo with the University of Guyana's Guyanese Languages Unit, then shipped IRIS, a Creolese-speaking assistant on WhatsApp. Their fine-tuned BART-Large writes Creolese nine times better than GPT-4 on the same test set. Devonish is also the linguist behind the Jamaican Language Unit's orthography work, which is worth registering on its own: the region has carried this across countries and decades on the shoulders of roughly the same small group of people.

Trinidad and Tobago · 2026
A small sample, expanded, then checked by hand

Nunes, Mohammed, Smith, Pooransingh and Ringis published a framework in the Caribbean Educational Research Journal for Trinidadian English Creole that starts from a small set of local samples, generates a larger training set from them, and requires manual verification before retraining. Translation accuracy on their evaluation set rose from a BLEU score of 59 percent to 75 percent, with responses returning in about a second.

Jamaican Patois · 2022
The first reasoning benchmark in any creole language

Ruth-Ann Armstrong, John Hewitt and Christopher Manning built JamPatoisNLI and presented it at EMNLP 2022. It tests whether a model can tell that one Patois sentence follows from another, which is harder than translation because it requires holding the grammar rather than matching the vocabulary. Fluent Patois speakers double-annotated a validation sample and agreed with each other 89 percent of the time, setting the human ceiling the models were measured against.

Orbital Brand Science and CAIA are building the missing layer

Orbital Brand Science is working with the Caribbean AI Association on a project to build the missing layer directly: a structured record of Caribbean dialects designed to do two jobs at once. The first is to give future models the training data they have never had for this region. The second is to make today's models usable in Caribbean speech now, without waiting for anyone else to retrain a foundation model on our behalf.

Separating those two jobs matters. Training data for future models is a corpus problem: hours of transcribed audio, aligned parallel text, one orthography per language, coverage across territories, registers and ages. Making current models work is an adaptation problem: fine-tunes on open models, example banks, orthography converters and retrieval layers sitting between a Caribbean user and a system never built for them. The first pays off over years. The second can pay off in weeks, and Exhibit 1 shows the size of the prize.

Editor's note before publishing: confirm and insert the project's current stage, launch date, phase-one territories, target hours per language, named leads at Orbital Brand Science and CAIA, how contributors take part, and any funding or institutional partners. This section is written to the level of detail supplied and makes no unverified claims about scope, timing or resourcing.

Whatever shape it takes, any Caribbean language corpus has to answer four questions before it records a single hour, and every one of them is a governance decision rather than a technical one.

Howard kept the dataset. Google got the use of it.

Howard University retained ownership and licensing of its African American English dataset and serves as steward for its use. Access went first to researchers inside the historically Black college and university network, with wider release deferred. Google got the use of the data to improve its products. Howard kept the asset. Any Caribbean equivalent that hands raw speech to a foreign lab under a permissive licence has funded someone else's model with the region's voice and received a press release for it.

One orthography per language, decided once

Cassidy-JLU for Jamaican. Cave-GLU for Creolese. The National Kriol Council standard for Belizean Kriol. The Akademi Kreyòl Ayisyen standard for Haitian. The Papiamento Academy Foundation standard for the ABC islands. The 24-letter Kwéyòl system for Saint Lucia and Dominica. Mixing systems teaches a model to spell inconsistently, which is exactly the failure GuyLingo reported.

Speakers and transcribers have a claim

Meta paid its field partners for the Omnilingual corpus and said so in the paper. Speakers, transcribers, artists and estates whose recordings become training data have a claim, and settling it at the start costs less than settling it in court later.

Consent has to name the permitted use

A grandmother who records forty minutes of storytelling to help preserve her language has not agreed to have her voice cloned for an advertisement. Consent has to name the permitted use, not just the transfer of the file.

The honest objection
There is a real argument against building this, and it deserves stating rather than burying
This Newsletter's position is that the corpus should be built. The case against it is not weak, and the region should go in with its eyes open.

The case for building itSilence is also a decision

  • The threshold is reachableTen hours of transcribed audio per language clears the bar Meta's own results identify. A university department can do that in a term.
  • The method is publishedGeorgetown and St Augustine have each published a working approach, and Guyana's numbers show a small fine-tune beating GPT-4 ninefold on generation.
  • Public services are being automated nowEvery ministry chatbot deployed before the data exists locks the exclusion in for its full contract term.
  • Nobody else will do itMeta's field collection went to Africa and South Asia. Google's Woolaroo update named no Caribbean creole. Nine of sixteen territories are on nobody's roadmap.

The case againstYou are building someone else's asset

  • It is exactly what foreign labs lackA clean Caribbean speech corpus, released openly, makes creoles work in everyone's products, including for firms that will never hire a Caribbean engineer.
  • Voice cloning gets easier tooThe same recordings that teach a model to understand Patois teach it to synthesize Patois, and the region has no legal test case on synthetic voice.
  • The Howard model is untested hereRetained ownership with staged access is the best template available. No court in Jamaica, Trinidad or Guyana has ruled on who owns a voice inside a training set.
  • Standardizing can flattenOne orthography per language is technically necessary and will privilege some varieties over others. That cost lands on whichever varieties lose the vote.

The objection this Newsletter cannot fully answer is the third one. Whoever moves first will set the region's precedent on voice ownership through whatever contract they happen to sign, and it will be a commercial document rather than a considered piece of law.

What to do
Eight actions, ordered by how quickly they can be started
ActionWhat it changesLevel
Record and transcribe ten hours in your language Crosses the threshold at which 95 percent of languages in Meta's evaluation reached usable transcription accuracy Easy
Publish in the standard orthography Makes text machine-readable and stops the spelling fragmentation that hurt GuyLingo's results Easy
Test AI tools in creole before buying them Turns a vendor claim into a measurement. Ask for accuracy in the writing direction, which is where systems fail Easy
Digitize what your language unit already has Belize's dictionary and grammar, Saint Lucia's Kwéyòl materials and Jamaica's JLU output are one step from being training data Medium
Release broadcast archives under a stated licence Radio stations, media houses and record labels hold decades of transcribable Caribbean speech that currently sits idle Medium
Build parallel text, not just monolingual audio Translation models need the same content in both languages. Government notices, health leaflets and school material are the obvious first set Medium
Fine-tune an open model on your own data Reproduces the Guyanese result, where a 400 million parameter model beat GPT-4 by a factor of nine on writing creole Advanced
Write creole performance clauses into government procurement Makes accuracy on national speech a contract term rather than a hope, for every public-facing system bought from here on Advanced
By role
Who should move first, and on what

Governments and ministries

Before signing any public-facing AI contract, require a measured accuracy figure on national speech rather than on standard English, and require it in the direction the system will actually generate. UNESCO's Caribbean AI Policy Roadmap, launched in 2024 and endorsed by CARICOM's COTED-ICT on 7 July 2026, gives the region a framework it has already agreed to. Language performance belongs inside it. Eric Falt, UNESCO's Regional Director, put the wider point at that meeting: the Caribbean should help shape how AI is governed rather than simply consume technologies developed elsewhere.

The nine unmapped territories

Barbados, The Bahamas, Saint Lucia, Dominica, Grenada, Saint Vincent and the Grenadines, Antigua and Barbuda, Belize and Suriname each need one thing to change status: ten hours of transcribed audio in a consistent orthography. Several already have the orthography. Belize has a dictionary and a grammar sitting ready. This is the cheapest cultural policy available to a small state right now.

Universities and language units

The orthographies, the grammars and the trained linguists exist across UWI's campuses, the University of Guyana, Saint Lucia's Folk Research Centre and Belize's National Kriol Council. What is missing is transcription capacity and a data-sharing agreement between them. Ten hours per language, transcribed to one standard, would put every English-lexicon Caribbean creole above the threshold inside an academic year.

Media houses, radio and record labels

The largest untapped store of Caribbean speech is broadcast and recorded music archives. Anyone holding decades of tape is holding a training corpus and, if the licensing is handled properly, an asset. Transcription cost is the barrier, and it falls every year.

Businesses deploying AI to customers

If your call centre, chatbot or voice system was tested only in standard English, it has not been tested. Run it against how your customers actually speak, and watch what it produces rather than what it appears to understand. A system can look competent on comprehension and be unusable on output, which is the gap Exhibit 1 measures.

Artists, writers and creators

Work published in a standard orthography enters the machine-readable record. Work published in ad hoc spelling largely does not. That is an unglamorous reason to use Cassidy-JLU, Cave-GLU or the Kriol Council standard, and it is currently the strongest one.

Frequently asked questions

It is doing both thinning and preserving at once, in different places. The thinning is measurable: tested on Guyanese Creolese, GPT-4 reads at 29.80 BLEU and writes at 1.35, so what it produces is a blended approximation rather than the language. The preserving is also measurable: the Kreyòl-MT project created the first parallel text that has ever existed for 21 creole languages in 2024. Which effect dominates in a given territory depends entirely on whether anyone local has built a dataset.
Partly, and only in one direction. Large models usually extract the rough meaning of written creole because it shares much of its vocabulary with English. Writing back is where they fail, which Meta's March 2026 research calls the generation bottleneck. In practice a model reads your Patois, catches the gist, and answers in English or in an unconvincing imitation.
Three creoles from this region. Haitian Creole has been supported for years. Jamaican Patois and Papiamento were both added in June 2024, when Google put 110 new languages into Translate and took its total to 243. Bajan, Bahamian Creole, Kwéyòl, Vincentian, Grenadian, Antiguan, Belizean Kriol and Sranan Tongo are not currently listed.
Because Jamaican Patois has far more digital text than any other English-lexicon Caribbean creole, thanks to music, film and a large diaspora that writes online. When a model meets Bajan, Vincentian or Antiguan speech, the nearest well-represented pattern it holds is Jamaican, so it pulls toward it. The result sounds broadly Caribbean and is wrong in the specific. Jennifer Carolina Gómez Menjívar and William Salmon documented the same crowding-out with no technology involved, in Tropical Tongues (University of North Carolina Press, 2018): Kriol rose in Belize after independence while Mopan and Garifuna weakened.
Ten hours for speech, and under two thousand sentence pairs for text. For speech, ten hours of transcribed audio is the practical floor: in Meta's Omnilingual ASR results from November 2025, 95 percent of languages with at least ten hours achieved a character error rate below 10, against only 36 percent of those with less. For text, the GuyLingo team fine-tuned models on fewer than 2,000 sentence pairs and produced output nine times better than GPT-4 in the writing direction. Neither is a national-scale undertaking.
Linguistically, Caribbean creoles such as Jamaican, Creolese, Kriol and Kwéyòl are languages in their own right, with grammars distinct from their European vocabulary sources. They emerged from contact between African and European languages during slavery. It matters technically because a model treating creole as misspelled English applies English grammar rules and fails on tense, aspect and pronoun systems that work differently. The measured proof is the direction gap: GPT-4 handles Creolese comfortably when the target is English and collapses to 1.35 BLEU when the target is Creolese.
Nobody in the region has tested this in court, which is the honest answer. The strongest available template is Howard University, which retained ownership and licensing of the African American English dataset it built with Google Research, released it first to historically Black colleges and universities, and deferred wider access. Google received use of the data while Howard kept the asset. Caribbean institutions building language corpora should assume they are setting regional precedent through whatever contract they sign, because at present they are.
Test the output, not the input. Write twenty sentences the way your customers actually speak, ask the tool to reply in the same register, and show the result to a fluent speaker. If a vendor cannot give you a measured accuracy figure for generating Caribbean speech rather than merely understanding it, they have not tested it either. For a government, procurement is the move that changes most: make creole accuracy a contract term. For a university or language unit, it is ten hours of transcribed audio in a consistent orthography.
Machines learn a language from whatever record of it exists. Where the record is thin, they invent, and the invention travels back to the speakers as a corrected version of themselves. Where the record is good, they hold the language steady for longer than any single generation can. Nine of our sixteen territories have not yet decided which of those they are building.
Caribbean AI Newsletter · The Culture Machine, 18 August 2026
🌴
About the Caribbean AI Newsletter

The Caribbean AI Newsletter is the leading daily source for artificial intelligence news, analysis, and practical insight for the Caribbean. The newsletter covers policy, workforce, cybersecurity, climate resilience, governance, language and cultural data, regional research, and the founders building the region's AI sector.

JamaicaTrinidad and TobagoBarbadosGuyanaThe BahamasSaint LuciaGrenadaSaint Vincent and the GrenadinesAntigua and BarbudaDominicaBelizeSurinameHaitiDominican RepublicCuracaoAruba
Visit the Directory
Next
Next

Pitchforks for AI Users: Invisible Watermarks in Claude AI