You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
fix: report the right line and column for CRLF and mixed line endings - #642
The line and column reported in a parse error depend on the document's line endings. Source._to_linecol splits the source with splitlines(), which drops the terminators, and then assumes every break costs exactly one character:
cur+=len(line) +1
A \r\n break costs two, so the running offset falls one character behind for every line before the error, and the error lands on the wrong line and the wrong column. exceptions.py documents that a ParseError "references the line and location within the line where the error was encountered", and toml_file.py:36-44 only normalises \r\n when the whole document uses it consistently — so a file with mixed endings read from disk reaches this path today.
The fix lets each split line carry its own terminator (splitlines(keepends=True)) and uses that length, so the offset matches the characters actually consumed. Documents that use only \n come out identical; CRLF and mixed documents now report the position the LF form already did.
Test: pytest tests/test_parser.py -k position fails on main (858186a) — the CRLF case reports UnexpectedCharError("Unexpected character: '@' at line 4 col 0") where the character is on line 3, column 4 — and passes on this branch. The full suite (pytest --ignore=tests/test_toml_tests.py -q) is 420 passed, and ruff check / ruff format --check at the version pinned in .pre-commit-config.yaml (0.16.10) are clean. I also walked every character index of every file under tests/examples and compared the reported position before and after: 6913 positions disagreed for CRLF copies of those documents and 0 disagree after the change, while LF-only documents show 0 differences either way.
Not run: tests/test_toml_tests.py, because the tests/toml-test submodule is empty in this checkout. I read it — it asserts TOMLKitError is raised for invalid documents and compares parsed JSON for valid ones, and does not look at line/column — but that is a reading, not a run.
Notes: diagnosis and patch drafted with the assistant; every number above comes from commands run on this checkout. AI-assisted development. Human review requested before merge.
The one failing check here is not from this change. Integration / Ubuntu / 3.12 / poetry fails inside Poetry's own suite, not tomlkit's:
FAILED tests/utils/test_helpers.py::test_extractall_sdist_no_symlink_path_traversal_via_directory_symlink
E Failed: DID NOT RAISE OutsideDestinationError
1 failed, 3309 passed, 34 skipped, 2 xfailed in 56.64s
The same job fails identically on other open PRs that do not touch parsing — run 38073613876 on fix-array-inplace-repeat-202 (2026-10-10 17:54) reports the same test with 1 failed, 3309 passed, and run 37986342804 (2026-10-09) also fails that job. master's own integration run passed on 2026-10-07, so this looks like Poetry's CI moving under the integration job rather than anything in this branch.
The Integration / Ubuntu / 3.12 / poetry failure is not from this change. It is poetry's own tests/utils/test_helpers.py::test_extractall_sdist_no_symlink_path_traversal_via_directory_symlink (Failed: DID NOT RAISE OutsideDestinationError), and that job checks out poetry's current main rather than a pinned tag, so it moves independently of tomlkit. Two other open PRs fail it identically with the same totals — run 38073613876 (head 0c09e04) and run 37986342804 (head b480695): 1 failed, 3309 passed, 34 skipped, 2 xfailed.
This branch's own Tests run succeeded on head 911db3d, and the diff is tomlkit/source.py plus tests/test_parser.py only — _to_linecol cannot reach tar extraction.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The line and column reported in a parse error depend on the document's line endings.
Source._to_linecolsplits the source withsplitlines(), which drops the terminators, and then assumes every break costs exactly one character:A
\r\nbreak costs two, so the running offset falls one character behind for every line before the error, and the error lands on the wrong line and the wrong column.exceptions.pydocuments that aParseError"references the line and location within the line where the error was encountered", andtoml_file.py:36-44only normalises\r\nwhen the whole document uses it consistently — so a file with mixed endings read from disk reaches this path today.The fix lets each split line carry its own terminator (
splitlines(keepends=True)) and uses that length, so the offset matches the characters actually consumed. Documents that use only\ncome out identical; CRLF and mixed documents now report the position the LF form already did.Test:
pytest tests/test_parser.py -k positionfails on main (858186a) — the CRLF case reportsUnexpectedCharError("Unexpected character: '@' at line 4 col 0")where the character is on line 3, column 4 — and passes on this branch. The full suite (pytest --ignore=tests/test_toml_tests.py -q) is 420 passed, andruff check/ruff format --checkat the version pinned in.pre-commit-config.yaml(0.16.10) are clean. I also walked every character index of every file undertests/examplesand compared the reported position before and after: 6913 positions disagreed for CRLF copies of those documents and 0 disagree after the change, while LF-only documents show 0 differences either way.Not run:
tests/test_toml_tests.py, because thetests/toml-testsubmodule is empty in this checkout. I read it — it assertsTOMLKitErroris raised for invalid documents and compares parsed JSON for valid ones, and does not look at line/column — but that is a reading, not a run.Agent Drafting Metadata