·

Decoding Animal Communication: A Better Way to See Hidden Risk

Machine learning is being pointed at whale song and ape gesture with real money behind it. A new Topoi paper by Marriah Alcantara and Kristin Andrews sets out the research ethics, and makes a claim that reaches well beyond biology: a system that genuinely interprets does not only reveal a moral situation, it creates one.

Decoding animal communication in five moves: structure is not meaning, interpretive cost was the privacy control, emitting is not reading, understanding creates duty, and Bruno asked to be let out.
In this article9 min read

Bruno the chimpanzee learned American Sign Language at the University of Oklahoma. When the funding ran out he was sent to a biomedical facility, where he had nobody to sign with. Years later one of the researchers from the old project came to visit, and Bruno began to sign. What he said was KEY OUT. He was asking to be let out, and that is the scene Marriah Alcantara and Kristin Andrews use to explain what is really at stake in decoding animal communication.

The hard part was never hearing him. It was what hearing him obliged, and the fact that nobody could meet it.

Their paper in Topoi, published in April 2026, argues that machine learning translation of animal signals is worth attempting, enumerates the ways it can go wrong, and then makes a claim that should interest anyone who builds interpretation systems for a living. A working translation would not only describe a moral situation. It would create one.

You do not get to build a system that understands something and then decide the understanding does not oblige you.

Decoding animal communication is funded engineering now, not a thought experiment

It is easy to file this under speculation. The funding says otherwise.

Alcantara and Andrews open their paper on three programmes: the Earth Species Project, Project CETI, which applies machine learning to sperm whale codas off Dominica, and the Interspecies Internet. Behind them sits a prize. The Jeremy Coller Foundation and Tel Aviv University run the Coller Dolittle Challenge, which offers “an annual prize of USD100,000” and, for the team that demonstrates genuine two-way interspecies communication, “either a USD10 million equity investment or a USD500,000 cash prize”.

$100,000

Awarded annually by the Coller Dolittle Challenge to work advancing two-way communication with another species.

$10 million

The major prize, as an equity investment, for a team that actually cracks it. A $500,000 cash alternative is offered instead.

Zero-shot

NatureLM-audio classifies species it was never trained on, prompted in ordinary language.

The technical state of the art is further along than the public conversation suggests. NatureLM-audio, from the Earth Species Project, is an audio-language foundation model trained on bioacoustic archives alongside human speech and music. Earth Species describes it as able to “solve a wide range of bioacoustics tasks zero-shot” from a natural language prompt: naming species, predicting life stage, identifying call types.

The biology gives the attempt a reason to exist. The paper surveys referential alarm calls in vervet monkeys, babbling in young songbirds, Zipf-like statistical distributions in humpback whale song, cultural transmission of that song between populations, and stable gestural meanings catalogued in the Great Ape Dictionary.

None of that is translation. Every one of those results is a pattern, and the authors are explicit that the models find patterns rather than confirmed meanings. That distinction is the whole governance problem in decoding animal communication, and it does not stay in the ocean.


A pattern you can reproduce is not a meaning you can defend

A model that reliably groups sperm whale codas has demonstrated that the codas have structure. It has not demonstrated what any of them are for.

Alcantara and Andrews call this opacity, and they compare it to algorithmic bias in domains where nobody could inspect the reasoning either. A traditional playback experiment uses a recording of a real animal saying a real thing. A generated signal carries no such guarantee, and the scientists running the experiment may have no way to establish what they have just said.

The failure mode is familiar to anyone who has watched a fluent summary win a meeting. Output quality is not evidence of comprehension, and a confident interpretation is harder to argue with than an uncertain one. That is the mechanism behind cognitive surrender to a persuasive answer, transposed to a species that cannot correct the record.

Here is a tension the paper’s account implies, and this reading is mine rather than theirs. Train on labels supplied by human experts and the model inherits what those experts already believed, which makes genuine discovery unlikely. Train without labels and the model may find something real that nobody can check. Neither branch gives you a meaning you could defend to a sceptical committee, and the second branch is the one that gets published.


Privacy has been protected by nobody bothering to interpret

The most transferable idea in the paper is also the quietest.

Decoding animal communication does not begin with access. Humans have always been able to hear animals, and whale song has been on record since the 1960s. What has protected animals is not inaudibility. It is that nobody could work out what any of it meant, and the authors argue that a working translation removes that barrier rather than crossing a new one.

They go further, and this is the part worth sitting with. Animals already behave as though information about them matters. The paper cites bird species that change their food-caching behaviour when potential thieves are watching, and “the use of vocal suppression to conceal courtship and mating behaviors in blackbirds, African lions, and chimpanzees”. These are strategies for managing who knows what. A translation system dismantles them without ever asking.

Now move it across. Most personal data in most organisations is not protected by access control. It is protected by the fact that interpreting it was too expensive to bother with. Call recordings, keystroke logs, badge data, the shape of somebody’s calendar: all lawfully held, most of it never read, and safe largely because reading it cost more than it was worth.

Interpretation is now cheap. The protection was never a control, and a data protection impact assessment written when the cost was high is describing a world that no longer exists. This is the practical reason frameworks such as the NIST trustworthy AI work keep insisting on documented purpose rather than documented permission.

The check is unglamorous and takes an afternoon. List the datasets your organisation holds but has never systematically read, then ask what a competent model would extract from each one this year. That list is your real exposure, and it is almost never the same as the list of systems already in scope.

Decoding animal communication risk ladder from record and classify through interpret to generate and broadcast, with the new risk that appears at each rung and the control it needs.
Each rung adds a risk the rung below did not have. Most governance is written for the first two and applied to all five.

Reading a signal and producing one are different acts

The paper is sharpest on what happens when a model stops listening and starts speaking.

Play a generated vocalisation into a wild population and several things can go wrong at once. The message may be unknown to the people broadcasting it, and may provoke a response that costs the animal energy it cannot spare. The signal may appear to come from an individual who does not exist, which is a social problem, not an acoustic one. Their comparison is the obvious one. This is a deepfake, aimed at a community with no way to check it.

Worse, some of these signals are culturally learned. Humpback song propagates between populations. A synthetic contribution does not simply fail. It can persist.

Any team that has moved a model from classification to generation knows the shape of this, even if the stakes were smaller. Reading is recoverable. Emitting is not always recoverable, and the controls that made the read side safe do not transfer. The discipline needed is the one set out in engineering around uncertainty rather than around confidence: constrain what the system is permitted to emit, not merely how well it scores.


Decoding animal communication would create obligations, not only reveal them

This is the argument that makes the paper worth an hour of a governance practitioner’s time.

Translating a language with no dictionary and no bilingual speaker is what Quine called radical translation. It runs on the principle of charity: you have to assume the speaker mostly believes true things and is acting rationally, because otherwise nothing they emit can be evidence of anything. Alcantara and Andrews press the point that this assumption is already built into the attempt. Every team pointing a model at animal signals is presupposing a rational agent at the other end, whether or not anyone says so out loud.

If the translation works, the presupposition was correct.

From there they draw on Stephen Darwall’s second-person standpoint. Two-way communication is not a data channel. It puts both parties into a relationship of mutual accountability, in the way that a promise or an apology creates a fact that was not there before. Their conclusion is that success would “reveal an existing moral landscape and create a new one”, and that the second half is the part nobody is funding.

Decoding animal communication and the second-person standpoint: before translation a signal is data with no duty attached, after translation it is a claim from a party you are now accountable to.
The system does not only change what you know. It changes who has standing to make a claim on you.

Which returns us to Bruno, and to the reason his two signs are the right frame rather than a sad anecdote. The communication worked. The obligation landed. Nobody at the institution could discharge it, and the gap between those two facts is the whole problem.

Substitute your own system. A model that reliably infers distress from a customer’s messages has created an obligation to act on distress. One that infers burnout from engineering telemetry has created an obligation to a named person. The capability arrives first and the duty arrives with it, uninvited, whether or not there is a process to absorb it. Deciding what you will do about the interpretation belongs at design time, alongside the question of who holds judgement when the model speaks first.


What the paper can and cannot tell us

This is a philosophy paper about decoding animal communication. It is an argument, not an experiment, and it reports no result of its own. Nothing in it demonstrates that any animal signal has been translated, because nothing has been.

The empirical work it surveys is real but contested. The position associated with Noam Chomsky, that animal communication differs from human language in kind rather than degree, remains a live one, and the paper acknowledges it rather than defeating it. A Zipf-like distribution in humpback song is a statistical signature that also appears in systems nobody would call linguistic. Compositional syntax in a handful of species is not a grammar.

Its central claim is conditional throughout. Successful two-way communication would entail rational agents in normative communities. Whether that success is reachable is exactly what remains open, and the authors do not pretend otherwise.

One small check while reading it. The paper describes the Coller Dolittle grand prize as ten million dollars in equity; the challenge’s own page offers either that equity investment or a $500,000 cash prize. The difference changes nothing in the argument, and it is a reminder to read the primary page rather than the citation.

The transfer to enterprise systems, the risk ladder, the reading of interpretive cost as an accidental privacy control, and the supervised and unsupervised tension are Beta Tester Life interpretations. Alcantara and Andrews write about animals. They have not endorsed any of this, and none of it is legal or data protection advice.

Sources used

  1. Marriah Alcantara and Kristin Andrews, Can we talk to the animals? The ethics of using machine learning to decode animal communication, Topoi, 4 April 2026, open access; DOI 10.1007/s11245-026-10409-2.
  2. Jeremy Coller Foundation and Tel Aviv University, The Coller Dolittle Prize for Two-way Inter-species Communication, source of the prize figures quoted above.
  3. David Robinson, Marius Miron, Masato Hagiwara, Benno Weck and colleagues, NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics, arXiv:2411.07186, November 2024.
  4. Earth Species Project, Introducing NatureLM-audio, 11 November 2024.
  5. Project CETI, the sperm whale communication programme referenced in the paper.

The interesting question in this research was never whether a machine can find structure in a whale’s noise. It will. Structure is what machines find.

The scarce capability is not the model. It is the willingness to be bound by what the model turns out to have heard, decided before the system is switched on rather than after somebody asks to be let out.

Kevin Campbell, writer behind Beta Tester Life

Behind the notebook

Written by Kevin Campbell

Thirty years of technology, delivery and organisational change—translated into practical thinking for people doing the work.

Continue the journey

One thought leads to another.

Scroll to explore

Conversation

Add to the thinking

Questions, experience and thoughtful disagreement are welcome.

Leave a Reply