【问题标题】:How does ls sort filenames?ls 如何对文件名进行排序?
【发布时间】:2017-03-04 09:36:39
【问题描述】:

我正在尝试编写一个模拟 Unix 中 ls 命令输出的函数。我最初试图使用 scandir 和 alphasort 执行此操作,这确实打印了目录中的文件,并且确实对它们进行了排序,但由于某种原因,这个排序列表似乎与文件名的相同“排序列表”不匹配ls 给出的。

例如,如果我有一个包含 file.c、FILE.c 和 ls.c 的目录。

ls 按顺序显示它们:file.c FILE.c ls.c 但是当我使用 alphasort/scandir 对其进行排序时,它会将它们排序为:FILE.c file.c ls.c

ls 如何对目录中的文件进行排序,从而给出如此不同的排序结果?

【问题讨论】:

  • ls 使用自己的自定义字符串比较,而alphasort 只是strcmp 的实现
  • @MDXF 我对 alphasort 有很多了解:/ 你知道 ls 的自定义字符串比较的细节,或者我可以在哪里阅读更多相关信息?我需要模仿 ls 的输出,所以我想了解它与 alphasort/strcmp 的不同之处。
  • ls 有一个基本的默认排序,它将事物按 strcmp 顺序排列,但它可能受语言环境的影响。要了解语言环境,您可以从 pubs.opengroup.org/onlinepubs/9699919799/basedefs/… 开始,特别是对于字符串比较,pubs.opengroup.org/onlinepubs/9699919799/functions/strcoll.html
  • source code 是您的最佳参考。
  • @kaylum 源代码如此庞大和复杂,几乎无法分辨是什么

标签: c linux bash sorting unix


【解决方案1】:

要模拟默认的 ls -1 行为,请通过调用使您的程序具有区域感知能力

setlocale(LC_ALL, "");

在main() 的开头附近,并使用

count = scandir(dir, &array, my_filter, alphasort);

其中my_filter() 是一个函数,它对以点. 开头的名称返回0,对所有其他名称返回1。 alphasort() 是一个使用语言环境排序规则的 POSIX 函数,与 strcoll() 的顺序相同。

基本实现类似于

#define  _POSIX_C_SOURCE 200809L
#define  _ATFILE_SOURCE
#include <stdlib.h>
#include <unistd.h>
#include <sys/types.h>
#include <sys/stat.h>
#include <fcntl.h>
#include <locale.h>
#include <string.h>
#include <dirent.h>
#include <stdio.h>
#include <errno.h>

static void my_print(const char *name, const struct stat *info)
{
    /* TODO: Better output; use info too, for 'ls -l' -style output? */
    printf("%s\n", name);
}

static int my_filter(const struct dirent *ent)
{
    /* Skip entries that begin with '.' */
    if (ent->d_name[0] == '.')
        return 0;

    /* Include all others */
    return 1;
}

static int my_ls(const char *dir)
{
    struct dirent **list = NULL;
    struct stat     info;
    DIR            *dirhandle;
    int             size, i, fd;

    size = scandir(dir, &list, my_filter, alphasort);
    if (size == -1) {
        const int cause = errno;

        /* Is dir not a directory, but a single entry perhaps? */
        if (cause == ENOTDIR && lstat(dir, &info) == 0) {
            my_print(dir, &info);
            return 0;
        }

        /* Print out the original error and fail. */
        fprintf(stderr, "%s: %s.\n", dir, strerror(cause));
        return -1;
    }

    /* We need the directory handle for fstatat(). */
    dirhandle = opendir(dir);
    if (!dirhandle) {
        /* Print a warning, but continue. */
        fprintf(stderr, "%s: %s\n", dir, strerror(errno));
        fd = AT_FDCWD;
    } else {
        fd = dirfd(dirhandle);
    }

    for (i = 0; i < size; i++) {
        struct dirent *ent = list[i];

        /* Try to get information on ent. If fails, clear the structure. */
        if (fstatat(fd, ent->d_name, &info, AT_SYMLINK_NOFOLLOW) == -1) {
            /* Print a warning about it. */
            fprintf(stderr, "%s: %s.\n", ent->d_name, strerror(errno));
            memset(&info, 0, sizeof info);
        }

        /* Describe 'ent'. */
        my_print(ent->d_name, &info);
    }

    /* Release the directory handle. */
    if (dirhandle)
        closedir(dirhandle);

    /* Discard list. */
    for (i = 0; i < size; i++)
        free(list[i]);
    free(list);

    return 0;
}

int main(int argc, char *argv[])
{
    int arg;

    setlocale(LC_ALL, "");

    if (argc > 1) {
        for (arg = 1; arg < argc; arg++) {
            if (my_ls(argv[arg])) {
                return EXIT_FAILURE;
            }
        }
    } else {
        if (my_ls(".")) {
            return EXIT_FAILURE;
        }
    }

    return EXIT_SUCCESS;
}

请注意,我故意使这比您严格需要的更复杂,因为我不希望您只是复制和粘贴代码。您可以更轻松地编译、运行和调查该程序,然后移植所需的更改——可能只是 setlocale("", LC_ALL); 行! -- 对你自己的程序,而不是尝试向你的老师/讲师/助教解释为什么代码看起来像是从其他地方逐字复制的。

上述代码甚至适用于命令行上指定的文件(cause == ENOTDIR 部分)。它还使用单个函数my_print(const char *name, const struct stat *info) 来打印每个目录条目;为此,它会为每个条目调用stat。

my_ls() 不是构造目录条目的路径并调用lstat(),而是打开目录句柄,并使用fstatat(descriptor, name, struct stat *, AT_SYMLINK_NOFOLLOW) 以与lstat() 基本相同的方式收集信息,但name是从descriptor 指定的目录开始的相对路径(dirfd(handle),如果handle 是打开的DIR *)。

确实,为每个目录条目调用其中一个 stat 函数是“慢”的(特别是如果您执行/bin/ls -1 样式输出)。但是,ls 的输出是供人类消费的;并且经常通过more 或less 让人们在闲暇时查看它。这就是为什么我个人认为“额外的” stat() 调用(即使不是真的需要)在这里是一个问题。我认识的大多数人类用户都倾向于使用ls -l 或(我最喜欢的)ls -laF --color=auto。 (auto 表示 ANSI 颜色仅在标准输出为终端时使用;即当 isatty(fileno(stdout)) == 1 时。)

换句话说,既然您有ls -1 订单,我建议您将输出修改为类似于ls -l(破折号,而不是破折号)。您只需要为此修改my_print()。

【讨论】:

  • 为了完整性和内容,你会得到一颗金星(嗯,我能做的最好的就是点赞:)
  • 你太棒了。非常感谢您在回答中超越自我!你不知道我多么感谢你的帮助。
【解决方案2】:

按字母数字(字典)顺序。

当然,这会随着语言而变化。试试:

$ LANG=C ls -1
FILE.c
file.c
ls.c

还有:

$ LANG=en_US.utf8 ls -1
file.c
FILE.c
ls.c

这与the "collating order" 有关。无论如何都不是一个简单的问题。

【讨论】:

  • 是的,它确实使用了(修改后的)strcoll。然而;从您的链接:“排序顺序是字典顺序”。
猜你喜欢
  • 1970-01-01
  • 2017-09-30
  • 2011-02-10
  • 1970-01-01
  • 2013-05-29
  • 2019-04-12
  • 1970-01-01
  • 2016-05-03
  • 2020-06-08
相关资源
最近更新 更多