Skip to content

Notebook · 2026-09-12 · updated 2026-10-10

How to fix CSV encoding problems

Tell UTF-8 from Windows-1252, recognize the replacement character, and re-read a CSV without damaging accents.

Encoding is the map from bytes in the file to characters on screen. If the map is wrong, accents, pounds, and euro signs break. The rows underneath are still there. You usually do not need to edit cells. You need to read the same bytes with the right encoding.

What the symptoms look like

  • café appears as café. The file is UTF-8 and something read it as Latin-1 or Windows-1252.
  • café appears as caf� or cafe?. The bytes were not valid in the encoding you chose, or a program already replaced them.
  • A blank box sits in front of the first header. That is often a byte-order mark made visible.

Which encoding to try

UTF-8 is the right default for new files. Excel on Windows has historically saved "CSV" as Windows-1252, which is close to Latin-1 but includes the euro sign. Older Mac files may be Mac Roman. UTF-16 shows up when a BOM of FF FE or FE FF is present. CleanCSV checks the BOM, then tries strict UTF-8, then offers Windows-1252. You can switch encoding in the file bar. That re-reads the original bytes. It does not guess by rewriting words.

If the replacement character is already stored in the file, no encoding can bring the original letter back. The information is gone. Go back to the system that exported the file and export UTF-8.

Saving so the next program agrees

When you download, you can add a UTF-8 BOM. Excel on Windows uses that mark to recognize UTF-8. The BOM is not part of the first cell. CleanCSV strips a BOM when it reads, so the header stays "Name" rather than a hidden character plus "Name". More Excel surprises are covered in Why Excel changes CSV values.

Open the cleaner