Skip to content

Tracking Issue for normalize_lexically #134694

Description

@ChrisDenton

View all comments

Feature gate: #![feature(normalize_lexically)]

This is a tracking issue for Path::normalize_lexically that is used to normalize a path, including .. parent references, without touching the filesystem.

Public API

// std::path

impl Path {
    pub fn normalize_lexically(&self) -> Result<PathBuf, NormalizeError> {
}

Steps / History

Unresolved Questions

  • Should PathBuf have a method that normalizes in-place? Normalizing a path just removes components.
  • Should this disallow resolving to the empty path or the single . path.
  • Should this only allow paths without a root or prefix? It may be surprising that base_path.join(user_path.normalize_lexically()?) can still escape the base path.

The last two can be summarised as: should this strictly require paths to be sub-paths?

Footnotes

  1. https://std-dev-guide.rust-lang.org/feature-lifecycle/stabilization.html ↩

Activity

  1. added
    C-tracking-issueCategory: An issue tracking the progress of sth. like the implementation of an RFC
    T-libs-api[DEPRECATED; DO NOT USE]
    on Dec 23, 2024
  2. added 2 commits that reference this issue on May 25, 2025
  3. added a commit that references this issue on May 26, 2025
  4. ChrisDenton commented on May 27, 2025

    @ChrisDenton
    MemberAuthor

    The basic implementation is now implemented in nightly if people would like to experiment.

  5. added a commit that references this issue on May 30, 2025
  6. clarfonthey commented on Aug 13, 2025

    @clarfonthey
    Contributor

    Should this take &mut self and normalize in-place? Normalizing a path just removes components.

    This could be done for PathBuf, but &mut Path cannot actually delete anything, since the resulting path has to be the same length. Maybe it could be done in a way similar to partition_dedup where all the components are moved to the end, but that would be confusing and a PathBuf-based method sounds much better API-wise.

  7. ChrisDenton commented on Aug 13, 2025

    @ChrisDenton
    MemberAuthor

    We could have a function both on Path and PathBuf. The only issue would be choosing names to distinguish them. normalize_lexically is already a bit of a mouthful (somewhat intentionally since it shouldn't be what people reach for by default).

  8. clarfonthey commented on Aug 13, 2025

    @clarfonthey
    Contributor

    How would you do a function on Path that removes components of the path in-place, given what I mentioned?

  9. ChrisDenton commented on Aug 13, 2025

    @ChrisDenton
    MemberAuthor

    No, I meant have a function on Path that returns a PathBuf and a separate function on PathBuf that does in-place normalisation.

  10. clarfonthey commented on Aug 13, 2025

    @clarfonthey
    Contributor

    Ah, yeah, that makes sense. I was mostly commenting on the note at the end that taking &mut self would only work on PathBuf, meaning that it would be worthwhile to have both.

  11. ChrisDenton commented on Aug 13, 2025

    @ChrisDenton
    MemberAuthor

    I've updated it to suggest adding a method to PathBuf instead. If anyone wants to implement that then feel free, I probably won't get to it for awhile.

  12. MikkelPaulson commented on Feb 11, 2026

    @MikkelPaulson
    Contributor

    I'm not sure that it makes sense for this operation to be fallible. Given an input of "abc/../../def", an output of "../def" seems perfectly reasonable. Even /../etc, which seems like it should be an error, is silently resolved to /etc on Linux.

    There is absolutely a strong use case for checking if a path escapes the pwd, but normalize_lexically:

    a) isn't a one-stop shop for this because it will happily represent absolute paths; and
    b) isn't an efficient solution to that problem because it allocates.

    With that in mind, I suggest always returning a PathBuf rather than Result<PathBuf, _>.

    On that note, does it make sense to return a Cow<'_, Path> to avoid allocating when no change is made? This is already done for eg. with_trailing_sep().

  13. clarfonthey commented on Feb 12, 2026

    @clarfonthey
    Contributor

    I definitely agree that it would be reasonable for this API to return Cow<'_, Path>, alongside my previous recommendation to add a version that just applies to PathBuf in-place.

    Also in agreement that it would be reasonable for the final result to contain relative paths, assuming that any .. are at the beginning of the path. I mean, after all, abc is a relative path.

  14. 39 remaining items

  15. eirnym commented on Aug 28, 2026

    @eirnym

    I also wore to always succeed if path can be opened. I'm not sure about enum as proposed above. This sounds more like "classify my path", which is not direct responsibility of this function.

    For example, relative paths "../" and "c:../dir" can be opened, so function must succeed.

  16. added a commit that references this issue on Aug 29, 2026
  17. added a commit that references this issue on Aug 29, 2026
  18. added a commit that references this issue on Aug 29, 2026
  19. schneems commented on Sep 10, 2026

    @schneems
    Contributor

    For example, relative paths "../" and "c:../dir" can be opened, so function must succeed.

    I think the limitation of the function (as implied by "lexically") is that it doesn't perform any disk lookups. i.e. it won't ever produce a std::io::Error. Like calling std::fs::absolute would if the call to CWD fails. Without knowing CWD, you cannot produce a meaningful normalization of ../ since we have no way to know what is being eaten/folded.

    To the concrete proposal: There are cases where we cannot produce a meaninfully different output than what we're given like ../ for those cases I would prefer to hint "nothing changed" or "could not normalize". Versus if we give them an enum, they have to do the mental math to figure out which variants hold normalized outputs and which don't. I think more than a few lazy programmers might just reach for the Into<PathBuf> and completely miss that their input couldn't be normalized. So, I prefer to differentiate them with an Err.

    Zooming out: I sat with the core of "what exactly is normalization" and wrote some thoughts on a few different ways to think of it, as well as a straw-man proposal. https://gist.github.com/schneems/e0f484bef06ba452199b3eed776203a3.

    Applying my proposal, I would suggest these for relative paths:

    • Empty input "" → "."
    • CurDir input "." → "."
    • Eat everything "a/.." → "."
    • Plain relative "a" → "./a"
    • Dot relative "./a" → "./a"
    • Starting dot slash ./a/b → ./a/b
    • Slash no dot a/b → ./a/b
    • Relative escape ”.." → Err
    • Relative escape ”../anything/else" → Err

    So a normalized relative path ALWAYS starts with a . and never ends with one (Path strips those already).

    Absolute edge cases:

    • Root "/" → "/"
    • Absolute eat everything "/a/.." → "/"
    • Absolute escape "/../.." → "/" (pin to root)
    • Absolute escape multiple "/../../../a" → "/a" (pin to root)

    I'm also now in Zulip, under "Richard Schneeman" if you want to peel off a part of that and have more of a chat.

  20. eirnym commented on Sep 11, 2026

    @eirnym

    There was no intention to make this API to touch filesystem at all.

    My point is that lexical normalisation is a part of link traversal, but not all of it. And it's better to move this responsibility to the caller as all situations might be different.

    E.g. .. in one context could mean "went beyond safety boundary, in other case it's just a path on a filesystem.

    If there's an intention to have such api, there's a method of Path to check if it's a relative path which fully covers the second safety part in their interpretation.

    The only path which is in a gray area is empty path which may or may not be considered as equivalent of ..

  21. schneems commented on Sep 11, 2026

    @schneems
    Contributor

    @eirnym i still guess I don't totally follow. If you want to represent the traversal there's std::fs::canonicalize.

    E.g. .. in one context could mean "went beyond safety boundary, in other case it's just a path on a filesystem.

    Really, the safety provided by lexical normailzation is for containment is limited as even a normalized path could have a symlink that points to a prior path. I still think it's valuable in an application to reduce/remove .. as it makes reasoning about the logic easier. It reduces the entropy. But it can also cut the other way i.e. a .. can "escape" a symlink such that folding it lexically does not produce an equivalent path as a physical traversal. Ruby's File.expand_path both follows physical links and folds as it goes. I think that's safer but I think there's some subtle dragons there too as you can wind up in situations where the traversal stops due to a problem like a broken symlink.

    Most languages seem to have a number of slightly different APIs with slightly different behavior. That's why I'm suggesting we try to lock down what "is" and isn't the job of this current feature. That way we can have strong confidence when a PR comes in to accept or reject it. And if someone needs a thing that isn't provided here, it would be a hint to start the process over with a new API with a new stated purpose and goal. I don't think we can have a single API make everyone 100% happy for all their needs.

    The only path which is in a gray area is empty path which may or may not be considered as equivalent of ..

    There's a few ambiguous cases in the existing implementation beyond this one (spelled out some differences in a few of my links). I think it's more useful to define this behavior and declare it produces something other programs can use.

  22. eirnym commented on Sep 11, 2026

    @eirnym

    The only think I'd like that lexical normalisation won't think for me in what context I'd like to use it. I'd like that relative paths were valid, even they follow above current folder.

  23. Skgland commented on Sep 11, 2026

    @Skgland
    Contributor

    While most of the choices seam reasonable/acceptable I have a few notes.

    Absolute Escape

    If lexical normalization should return an Err for any path it should be ̀the absolute root escapes, as that can't be check easily before or after.

    I.e.

    • Absolute escape "/../.." → "/" (pin to root)
    • Absolute escape multiple "/../../../a" → "/a" (pin to root)

    should result in Err.

    Relative Escape

    For me a relative path can fail to be normalized .. is perfectly fine at the beginning of a normalized path.
    So rather than

    So a normalized relative path ALWAYS starts with a . and …

    A normalized relative path either starts with either a "." or at least one ".." .
    "." may only appear as the first component (after an optional prefix)
    ".." may only appear as the first component (after an optional prefix) or after itself

    I.e.

    Relative escape ”.." → Err
    Relative escape ”../anything/else" → Err

    shouldn't result in Err. If someone wants to reject these checking for .. after normalization can be done easily.
    Though if you insist that these should Err as long as the Err provides the what you call a partially normalized path I could live with that.

    Plain relative

    I don't think it is good that we deviate here from all other referenced implementation

    • Plain relative "a" → "./a"

    Though I don't feel strongly about it as it should be easy to check for so that one can skip normalization in that case.

    Paths with Prefix

    The examples are missing e.g. windows paths with prefix.
    Basically I think it should act as
    1. if the path has a prefix remove it
    2. perform normalization as described before
    3. if 1. removed a prefix add it back

  24. ChrisDenton commented on Sep 11, 2026

    @ChrisDenton
    MemberAuthor

    I've started a discussion with the libs team on zulip. I'd like to keep that discussion focused on the broad objective of this method rather than the specific details. Without some agreement on what normalize_lexically is for we're unlikely to make progress.

  25. clarfonthey commented on Sep 13, 2026

    @clarfonthey
    Contributor

    Nomination for meeting discussion, with nomination shifting from #159863 to this.

  26. clarfonthey commented on Sep 22, 2026

    @clarfonthey
    Contributor

    So, collecting feedback from libs meeting:

    • We are in agreement that this should ignore symlinks, and that a/b/.. is always a.
    • We are in agreement that there are no parents of the root, so, /.. is just /. (+ for other prefixes as well, noting that C:.. is not C:\)
    • We know that this is not enough for security, and that you should not rely solely on this for security reasons. We could potentially find a resource to link here detailing some of the potential issues and why this isn't enough.

    We haven't figured out trailing slashes, yet, and it feels like this might be a place where we want to have extra options, or just not trim the trailing separator and let users do it themselves.

  27. ChrisDenton commented on Sep 22, 2026

    @ChrisDenton
    MemberAuthor

    Just to be explicit, the feeling in the libs meeting is that this function should just normalize the path and not do anything else, like restrict it to the root. E.g. a/../.. should be ../.

    As clarafon said, security of paths is a much more complex issue and we can't handle them all in one function. It's also not specific to lexical normalization. This kind of normalization may be one component of securing paths but it may also not be. Even in the case that it is part of securing paths, this one function can't do everything people want without growing a lot of options so it's better for it to do one thing so it can then be composable.

  28. removed
    I-libs-nominatedNominated for discussion during a libs team meeting.
    on Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    C-tracking-issueCategory: An issue tracking the progress of sth. like the implementation of an RFCT-libsRelevant to the library team, which will review and decide on the PR/issue.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions