Duplicate File Groups
Group duplicate repository files by content with deterministic output and path validation.
Implement find_duplicate_files(files), a repository scan helper that groups files with identical content.
Requirements
- Input is a dictionary from path to text content.
- Return only duplicate groups with at least two paths.
- Sort paths inside each group.
- Sort groups by their first path.
- Raise
ValueErrorinstead of skipping any invalid path: empty path; any path segment equal to"",".", or".."after splitting on/(covers empty path,.,.., repeated slashes, and trailing slashes such as"",../secret,src//app.py,src/./app.py,src/). - Different paths are distinct files; group by exact content equality.
- Empty file content is valid and can form a duplicate group; only invalid paths are rejected.
Example
Two paths have the same complete text. The third path is unique and disappears from the result. Empty text still counts as content:
1files = {"a.py": "x", "b.py": "x", "c.py": "y"}
2assert find_duplicate_files(files) == [["a.py", "b.py"]]
3assert find_duplicate_files({"empty-a.py": "", "empty-b.py": ""}) == [
4 ["empty-a.py", "empty-b.py"]
5]Sorting applies to paths, not to content or insertion order. For {"z": "x", "b": "y", "a": "y", "c": "x"}, first form the groups ["z", "c"] and ["b", "a"]. Sorting their paths gives ["c", "z"] and ["a", "b"]; sorting by the first path then returns [["a", "b"], ["c", "z"]].
Content identity and path identity
Compare the exact strings. A final newline changes content, so "x" and "x\n" aren't duplicates. Validation is independent of duplicate detection: {"only//file": "unique"} must raise ValueError even though it can't form a duplicate group. The function doesn't read a filesystem, resolve symlinks, or repair invalid path segments.
Constraints
- Keep the scan in memory.
- Work only from the input map.
- Make output deterministic.