【问题标题】:How to store frequently used data or parameters within an R package?如何在 R 包中存储常用数据或参数?
【发布时间】:2017-11-26 18:40:41
【问题描述】:

我正在创作一个 R 包,并且有几个数字向量,用户经常将其用作各种包函数的参数。将这些向量存储在包中以便用户轻松访问它们的最佳方式是什么?

我的一个想法是将每个向量保存为 inst/data 中的数据文件。然后用户可以在需要时使用数据文件的名称代替向量(至少,我可以在开发过程中这样做)。我喜欢这个想法,但不确定这个解决方案是否会违反 CRAN 规则/规范或导致任何问题。

# To create one such vector as a data file
octants <- c(90, 135, 180, 225, 270, 315, 360, 45)
devtools::use_data(octants)
# To access this vector in usage
my_function(data, octants)

我的另一个想法是创建一个单独的函数来返回所需的向量。然后用户将能够在需要时调用适当的函数。出于某种原因,这可能比数据更好,但我担心用户会忘记函数名称后的()

# To create the vector within a function
octants <- function() c(90, 135, 180, 225, 270, 315, 360, 45}
# To access this vector in usage
my_function(data, octants()) # works
my_function(data, octants) # doesn't work

有人知道哪种解决方案更可取或有更好的替代方案吗?

【问题讨论】:

    标签: r r-package


    【解决方案1】:

    说实话,我花了很长时间仔细阅读手册,问自己同样的问题。 去做吧,这是个好主意,很有用,而且有一些工具可以帮助你Writing help extension 手册描述了您可以保存数据的格式,以及如何遵循 R 标准。

    我建议在包中提供数据是使用:

    devtools::use_data(...,internal=FALSE,overwrite=TRUE)
    

    其中... 是您要保存的数据集的不带引号的名称。

    https://www.rdocumentation.org/packages/devtools/versions/1.13.3/topics/use_data

    您只需在包的inst 子目录中创建一个文件来创建数据集。我自己的例子是https://github.com/cran/stacomiR/blob/master/inst/config/generate_data.R

    例如我用它来创建 r_mig 数据集

    #################################
    # generates dataset for report_mig
    # from the vertical slot fishway located at the estuary of the Vilaine (Brittany)
    # Taxa Liza Ramada (Thinlip grey mullet) in 2015
    ##################################
    
    #{ here some stuff necessary to generate this dataset from my package
    # and database}
    setwd("C:/workspace/stacomir/pkg/stacomir")
    devtools::use_data(r_mig,internal=FALSE,overwrite=TRUE)
    

    这将以适当的格式保存您的数据集。使用internal = FALSE 允许所有使用data() 的用户访问。我建议您阅读data() 帮助文件。您可以使用data() 访问您的文件,包括当您不在包中时,只要它们位于数据子目录中。

    如果 lib.loc 和 package 都为 NULL(默认),则数据集为 在所有当前加载的包中搜索,然后在“数据”中搜索 当前工作目录的目录(如果有的话)。

    如果您使用的是 Roxygen,请创建一个名为 data.R 的 R 文件,其中存储所有数据集的描述。下面是 stacomIR 包中一个数据集的 Roxygen 命名示例。

    #' Video counting of thin lipped mullet (Liza ramada) in 2015 in the Vilaine (France)
    #' 
    #' This dataset corresponds to the data collected at the vertical slot fishway
    #' in 2015, video recording of the thin lipped mullet Liza ramada migration
    #'
    #' @format An object of class report_mig with 8 slots:
    #' \describe{
    #'   \item{dc}{the \code{ref_dc} object with 4 slots filled with data corresponding to the iav postgres schema}
    #'   \item{taxa}{the \code{ref_taxa} the taxa selected}
    #'   \item{stage}{the \code{ref_stage} the stage selected}
    #'   \item{timestep}{the \code{ref_timestep_daily} calculated for all 2015}
    #'   \item{data}{ A dataframe with 10304 rows and 11 variables
    #'          \describe{
    #'              \item{ope_identifiant}{operation id}
    #'              \item{lot_identifiant}{sample id}
    #'              \item{lot_identifiant}{sample id}
    #'              \item{ope_dic_identifiant}{dc id}
    #'              \item{lot_tax_code}{species id}
    #'              \item{lot_std_code}{stage id}
    #'              \item{value}{the value}
    #'              \item{type_de_quantite}{either effectif (number) or poids (weights)}
    #'              \item{lot_dev_code}{destination of the fishes}
    #'              \item{lot_methode_obtention}{method of data collection, measured, calculated...} 
    #'              }
    #'   }
    #'   \item{coef_conversion}{A data frame with 0 observations : no quantity are reported for video recording of mullets, only numbers}
    #'   \item{time.sequence}{A time sequence generated for the report, used internally}
    #' }
    #' @keywords data
    "r_mig"
    

    完整的文件在那里:

    https://github.com/cran/stacomiR/blob/master/R/data.R

    另一个例子:阅读:http://r-pkgs.had.co.nz/data.html#documenting-data

    然后你可以通过调用data("r_mig")在下面的测试中使用这些数据

    test_that("Summary method works",
        {
         ... #some other code
    
          data("r_mig")
          r_mig<-calcule(r_mig,silent=TRUE)
          summary(r_mig,silent=TRUE)
          rm(list=ls(envir=envir_stacomi),envir=envir_stacomi)
        })
    

    最重要的是,您可以使用手册中的内容来描述如何使用包中的函数。

    【讨论】:

    • 这很有帮助,谢谢!我仍然对r_mig 何时访问数据以及何时需要data("r_mig") 感到有些困惑。这是由data_use()internal 参数决定的,还是r_mig 仅在开发过程中可能,用户将始终需要data()
    • 你总是需要 data()。当你在你的包中编程时,你有一个数据目录,所以你可以很容易地访问你的数据。您需要做的是运行一次创建数据的文件(使用 use_data()),然后它们将可以在您的数据目录中访问
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-03-10
    • 2015-06-23
    • 1970-01-01
    • 2014-12-08
    • 2022-11-23
    • 1970-01-01
    相关资源
    最近更新 更多