When I prompt an LLM to write a four-part chorale in the style of Bach, the results are often unimpressive. Sometimes the LilyPond code contains errors. Often the part-writing is absurd.
But today I gave Claude Fable 5 my usual prompt, and the results were noticeably better than usual:

The writing is not perfect — there are still elementary mistakes, such as the parallel octaves in measures 10 and 13 — but it is the best I have seen so far. Compared to what LLMs were capable of two years ago, this is a major improvement.
Other LLMs
To get an idea of how good Fable 5 is, let’s compare it to two other models: GPT 5.5 and Gemini 3.5 Flash.
In response to the same prompt, GPT 5.5 exhibits two common errors — long sequences of parallel intervals and voices straying well outside their range:
Gemini 3.5 Flash, meanwhile, avoids these pitfalls,1 but fails in other ways, namely harmony. Most egregiously, it places a B-natural and a B-flat in the same chord. Even if we take one of these notes to be a typo, neither of the possible chords — B diminished or B-flat major — makes much sense.
In short, Claude Fable 5 is a big leap over its predecessors. It is the first LLM I have tested that has come anywhere close to passing the Bach benchmark.
Unlike Claude and GPT, Gemini did not one-shot the prompt. Its initial attempt contained LilyPond errors, which had to be fixed with a second prompt.




Is this benchmark "manual" - where you just review by hand weather the composition abides by musical theory and displays esteathic qualities similar to J.S. Bach?
Hi... I assume. you are familiar with David Copes work from the 90ties (Experiments in Musical Intelligence - EMI) and Douglas Hofsteaders "Turing test" for composition? https://computerhistory.org/blog/algorithmic-music-david-cope-and-emi/