Skip to content

Notebook · 2026-09-08 · updated 2026-10-10

How to remove duplicate rows in a CSV

Remove exact duplicate CSV rows while keeping the first or last copy, and avoid treating similar rows as the same.

A duplicate row is a later row whose every cell matches an earlier row. Same person with a different email is not a duplicate. Same email with a trailing space is not a duplicate either, until you decide that space does not matter.

Decide what "same" means

  • Exact match compares the full row, including letter case and spaces.
  • Trim outer whitespace first if " London" and "London" should match.
  • Do not convert IDs to numbers first. 00120 and 120 are different text values, and collapsing them destroys leading zeros.
  • Empty columns do not create duplicates by themselves, but a blank column that differs between otherwise identical rows will keep both rows.

Keep the first copy or the last

If the file is an export that appends corrections, the last copy is often the one you want. If the first row is the original record and later rows are accidental repeats, keep the first. Remove duplicate rows shows how many extra copies will go away and a sample of those rows. Unique rows stay. Nothing is deleted until you apply the preview.

Name,Amount
Ada,00120
Ada,00120
Ada,120

Keep first → Ada,00120 and Ada,120 both remain.
The second 00120 row is the only removal.

Check the count before you trust it

If the preview says thousands of rows will be removed and you expected dozens, stop. A shifted delimiter can make every row look alike or make real copies look different. Undo is available after you apply, and Revert to opened file puts the original grid back.

Open the cleaner