Support independent writingAbout the author →
Neil Meyer

My Music Training Dataset

Neil Meyer

An old Excel quiz, twenty badly pixelated album covers, an AI that did rather worse than expected, and a small experiment in what we bring with us when we recognise a pattern.

An old Excel quiz, twenty badly pixelated album covers, an AI that did rather worse than expected, and a small experiment in what we bring with us when we recognise a pattern.

Back in the early 2000s, when Excel macros were establishing their heavyweight role in everyday office life, I remember a spreadsheet being passed around that probably represented the pinnacle of the technology. It contained a collection of tiny album covers with the words removed, and below each one were boxes where you could type the artist and album name. You got a point for each correct answer.

It should be noted that 'Pink Floyd' and 'pink floyd' were different answers, and only one would give you the required point. Apparently, the future of automated assessment had arrived.

I had not thought about that spreadsheet for years, but it came back to me recently and I decided it might be fun to recreate something similar. The technical part was easy enough: find some memorable album covers, remove most of the useful information, reduce them until the people became arrangements of coloured squares, and put the results into a grid.

The harder part was deciding which albums belonged in it.

Twenty badly pixelated album covers arranged in a grid
Twenty badly pixelated album covers arranged in a grid

Before reading much further, have a look and see how many you recognise. Some may be immediately obvious. Others may sit irritatingly close to recognition without ever quite resolving into an artist or album. A few may mean absolutely nothing to you.

Then ask a slightly different question: what is missing?

I suspect that answer may eventually be more interesting than the score.

A fairly personal selection

My first pass produced sixteen covers, and at first they seemed like a reasonable collection of memorable albums. Looking at them together, though, I realised I had leaned quite heavily towards the 'popular' without considering enough of the 'personal'.

There was a lot of British and American rock and pop. Men were doing extremely well. Country music had apparently failed the admissions process, and the geographical centre of gravity had not travelled particularly far.

So I added four more.

Those additions were more deliberate. They broadened the collection, but they also made it more recognisably mine. Two came from the musical landscape I grew up with in South Africa, rather than simply from the things that had been most commercially prominent in Britain or the United States. Another moved the set firmly into country, and another addressed an increasingly obvious gap in female representation.

That created an interesting little problem of its own. The first sixteen were fairly close to an unprompted sample of whatever my memory supplied. The final four were editorial interventions made after I had inspected that sample and noticed some of its biases.

Was the final set broader? Yes.

Was it a slightly less pure representation of what first came to mind? Also yes.

If I had been trying to construct a balanced history of popular music, twenty albums would have been laughably inadequate anyway. I could keep adding more women, more Black artists, more countries, more genres and more recent music. Eventually I might create something respectable enough to survive a committee meeting, at which point it would probably tell us very little about the music that actually shaped me.

I'm sharing this because the grid is not intended to be definitive. It is a small part of my own musical training data.

And that training data has a noticeable cutoff.

When music was something you handled

For me, music was much more tactile through the 80s, 90s and into the early 2000s.

There was radio and MTV playing in the background, with somebody else largely curating what arrived, but a substantial part of listening was intentional and physical. If you wanted to hear something specific, you would thumb through a collection of records, cassettes or CDs, choose one, take out the physical media and engage with whatever machine was required to make it produce music.

Whilst the whirring of a CD drive spun up, or the crackling of a record announced the dropped needle, you could sit back with the album sleeve. Sometimes there were lyrics to read along with, sometimes photographs, production notes or fragments of the band's story, and sometimes there was simply artwork you stared at whilst the music played.

And, back then, you had paid for the whole album. Quite often, once you put it on, you listened to the whole album as well.

The physical object, the artwork and the music became linked in the experience. You saw the cover when you bought it, when you found it on the shelf, when you put it into the player and often whilst you were listening. It received repeated exposure simply because it occupied the same physical space as the music.

Spotify does not offer that same experience to me.

Don't get me wrong. I have discovered some wonderful new music through streaming, across both popular and niche artists, but I could not pick out the album cover for many of them. In quite a few cases, I am not convinced I could identify the artist from a line-up either. Instead, Discover Weekly serves up something interesting, I occasionally click the little 'like' button, and the song finds its way into a collection I may replay later.

The music still gets in. The visual identity often arrives with considerably less weight.

That probably explains why the albums that came readily to mind for this exercise started thinning out around the point at which my own listening shifted from physical ownership towards streaming. I did not stop discovering music; the form of the input changed.

That is a personal observation rather than a claim about everybody who uses Spotify. Somebody who has grown up with streaming may have an entirely different visual map of music, built around thumbnails, performers, playlists, videos and interfaces that barely existed when I was forming mine.

But my training is not your training, and it is certainly not AI's training.

Then I let GPT have a go

Once I had the full set of twenty covers, I thought it would be interesting to see how well GPT could identify them.

After all, it had been trained on a vast amount of information, several of the albums were hardly obscure, and it already had some contextual information about me from previous conversations and preset material.

How wrong could it be?

Well, as it happens, pretty wrong.

It correctly identified nine of the twenty.

The failures were more interesting than a random collection of guesses would have been. In many cases the answer was plausible. The colours, broad shapes or remaining fragments of composition had triggered an association with another well-known artist or album, and the model confidently completed the rest.

I cannot inspect the exact internal route by which it reached each answer, so I should not pretend I know that it 'preferred' something more popular or consciously decided that a relatively niche artist would be unlikely to appear. What I can observe is that, faced with incomplete visual information, it matched what survived against patterns already available to it, sometimes successfully and sometimes not.

That is also what makes the quiz possible for us.

If a handful of coloured shapes are enough for you to recognise an album immediately, the pixels themselves are not supplying everything you know. You are contributing some of the meaning from what you have seen before.

In that sense, prior exposure is not automatically a problem. It is also what makes expertise useful. An experienced doctor, engineer, mechanic or delivery leader will often notice patterns in incomplete evidence that somebody encountering the same situation for the first time does not.

The difficulty comes when a familiar pattern is the wrong one, and confidence arrives before verification.

The album quiz provides the harmless version of that problem. Nobody's life is materially affected because an AI confidently identifies the wrong record. The same mechanism becomes rather more important when models are asked to make judgements about people, roles, risk, suitability or behaviour.

Take the familiar example of asking an image-generation system to produce an executive. If the material it encountered during training disproportionately associated 'executive' with middle-aged white men wearing suits, that pattern can influence what appears in the output. Nobody needs to insert an explicit rule about what an executive should look like. Repeated representation can help establish the expected pattern.

Humans are not exempt from this. Our own internal datasets contain parents, teachers, books, television, music, workplaces, countries, cultures, friends, success, failure and a fairly random collection of things that happened to arrive at the right moment and stick. Some of those associations are useful, some are harmless, and some become problematic when we mistake familiar for normal.

That was already enough to make the experiment interesting. Then GPT made it better.

Marking its own homework

Later in the conversation, the AI discussed how well it had done on the test and told me it had correctly identified sixteen of the twenty covers.

It had not.

I went back through its answers and checked them. Nine out of twenty. Forty-five per cent.

It accepted the correction perfectly happily and produced a much better analysis afterwards, but the interesting part was not the arithmetic error itself. It had constructed a confident interpretation of a favourable result without first verifying whether the result was true.

There is something pleasingly human about remembering an examination rather more generously than the marking justified, but there is a more important point underneath it. The first failure had been in interpreting the input; the second came when evaluating the quality of its own output.

The second answer actually sounded more analytical. It was also more wrong.

That seems worth remembering when we discuss AI performance. A fluent explanation can arrive before the evidence underneath it has been checked, and once the explanation sounds coherent it can make the underlying claim feel rather more secure than it deserves.

It also exposes another layer of bias in my little experiment.

I chose the albums. I decided how much information to remove. I set the test. GPT interpreted the images. Then GPT briefly evaluated its own performance incorrectly.

At every stage, somebody or something had made choices about what mattered.

Even the pixelation itself was not as neutral as I initially assumed. Some album covers survive brutal reduction remarkably well because their identity rests on simple geometry, strong blocks of colour or a very familiar composition. Others depend on faces, text and fine detail that disappear almost immediately. Applying the same transformation to every image did not make the resulting challenge equally difficult.

That is a trivial observation when the task is identifying records. It becomes a more useful one when thinking about AI systems, because input data is rarely simply 'there'. Somebody selected it, cleaned it, labelled it, cropped it, transformed it, included some things and excluded others before a model ever had a chance to work with it.

The benchmark does not arrive from nowhere either.

If somebody roughly my age and with a similar background recognises seventeen of these covers whilst somebody much younger gets four, it would be ridiculous to conclude that the first person knows four times as much about music. I constructed an examination whose questions overlap heavily with my own cultural history.

Reverse the exercise and my apparent expertise could disappear very quickly. Give me twenty images drawn from current gaming culture, online creators, recent album artwork and whatever set of memes currently makes complete sense to a 16-year-old, and I suspect my responses would become increasingly optimistic guesses.

Benchmarks can still be useful. The important thing is to remember what they actually measured, who constructed them, and which assumptions were built in before the first score appeared.

A number does not remove those choices. Sometimes it just makes them harder to see.

Corrupting your training data

I also created a Spotify playlist containing a song from each of the twenty albums.

If you still want to identify the covers unaided, you may want to resist opening it for a while. Spotify displays artists, titles and artwork with very little respect for the integrity of my experiment.

Neil Meyer - Music Training Dataset on Spotify

It is not cheating if you are having fun.

The playlist also gives the experiment one final twist. If one of those artists means absolutely nothing to you today, perhaps you listen and discover something new. Maybe you recognise a song but had never registered the album it came from. Perhaps something that was part of my South African musical landscape has simply never crossed yours.

If I show you the same grid again in a month, there is at least a possibility that one of those previously meaningless patches of pixels will now resolve into something familiar.

I will, in a very small way, have corrupted your dataset.

I am deliberately using 'training' rather loosely here. Humans are not large language models, and pushing the comparison too far would turn a useful analogy into a bad explanation of both.

What interests me is the narrower point. What we encounter before affects how we interpret what arrives next. Sometimes prior exposure lets us recognise a pattern almost instantly. Sometimes it allows genuine expertise to operate on incomplete evidence. Sometimes the wrong familiar pattern arrives first and we confidently complete the picture incorrectly.

The model brings that history to the interaction.

So do we.

I started this by trying to recreate an old Excel music quiz and ended up with a rather more revealing picture of what happened to be stored in my own cultural memory. My parents are somewhere in the selection, as are South Africa, Britain, record shops, CDs, MTV, streaming and a collection of things I once thought sufficiently important to buy and keep on a shelf.

The omissions are in there too, precisely because they are missing.

So, assuming you have not cheated yet, how many of the covers could you identify? Which ones sent you confidently in the wrong direction? And, perhaps more interestingly, what album can you still not quite believe I left out?

Your answer to that last question probably says something about your training data as well.

I work with organisations navigating this shift, fractionally, as an adviser, or as a trusted collaborator. See how I work →

AI & Technology
← Back to all articles