排序包含语义版本的字符向量 [英] Sorting character vector containing semantic versions

查看:69
本文介绍了排序包含语义版本的字符向量的处理方法,对大家解决问题具有一定的参考价值,需要的朋友们下面随着小编来一起学习吧!

问题描述

似乎是一个非常基本的问题,但我无法真正找到一种简单"的方法.

Seems like a pretty basic question, but I can't really figure out an "easy" way to do it.

我想对character向量进行排序,该向量包含具有 base R功能的语义版本号 :

I'd like to sort a character vector containing semantic version numbers with base R functionality:

vsns  <- c("1", "10", "1.1", "1.10", "1.2", "1.1.1", 
           "1.1.10", "1.1.2", "1.1.1.1", "1.1.1.10", "1.1.1.2")

排序后应如下图所示:

# [1] "1"        "1.1"      "1.1.1"    "1.1.1.1"  "1.1.1.2"  "1.1.1.10"
# [7] "1.1.2"    "1.1.10"   "1.2"      "1.10"     "10"    

这并不能满足我的需求,因为R只是按字母顺序对整个内容进行排序:

This doesn't get me what I want, of cours, as R simply sorts the whole thing alphabetically:

sort(vsns)
# [1] "1"        "1.1"      "1.1.1"    "1.1.1.1"  "1.1.1.10" "1.1.1.2"  "1.1.10"  
# [8] "1.1.2"    "1.10"     "1.2"      "10"    
vsns[order(vsns)]
# [1] "1"        "1.1"      "1.1.1"    "1.1.1.1"  "1.1.1.10" "1.1.1.2"  "1.1.10"  
# [8] "1.1.2"    "1.10"     "1.2"      "10"    

尝试对其进行规范化(类似

Trying normalizing it (somewhat along this post), but I can't think of a matching/substitution scheme that would fit the structure of semantic versions:

tmp <- gsub("\\.", "", vsns)
# [1] "011"  "021"  "0101" "0201"
tmp_nchar <- sapply(tmp, nchar)
to_add <- max(tmp_nchar) - tmp_nchar
tmp <- sapply(1:length(tmp), function(ii) {
  paste0(tmp[ii], paste(rep("A", to_add[ii]), collapse = ""))
})
# [1] "10"       "1.10"     "1.1.10"   "1.1.1.10" "1.1.1.1"  "1.1.1.2"  "1.1.1"   
# [8] "1.1.2"    "1.1"      "1.2"      "1"   
vsns[order(tmp)]
#  [1] "1AAAA" "10AAA" "11AAA" "110AA" "12AAA" "111AA" "1110A" "112AA" "1111A" "11110"
# [11] "1112A"

到目前为止,我能想到的最好的方法是,但这似乎很...涉及;-)

The best I could come up with so far is this, but it seems pretty... Involved ;-)

sortVersionNumbers <- function(x, decreasing = FALSE) {
  tmp <- strsplit(x, split = "\\.")  
  tmp_l <- sapply(tmp, length)  
  idx_max <- which.max(tmp_l)[1]
  tmp_l_max <- tmp_l[idx_max]
  tmp_n <- lapply(tmp, function(ii) {
    ii_l <- length(ii)
    if (ii_l < tmp_l_max) {
      c(ii, rep(NA, (tmp_l_max - ii_l)))
    } else {
      ii
    }
  })
  tmp <- matrix(as.numeric(unlist(tmp_n)), nrow = length(tmp_n), byrow = TRUE)
  tmp_cols <- ncol(tmp)
  expr <- paste0("order(", paste(paste0("tmp[,", 1:tmp_cols, "]"), 
    collapse = ", "), ", na.last = FALSE",
    ifelse(decreasing, ", decreasing = FALSE)", ")"))
  idx <- eval(parse(text = expr))
  tmp_2 <- tmp[idx,]  
  sapply(1:nrow(tmp_2), function(ii) {
    paste(na.omit(tmp_2[ii,]), collapse = ".")
  })
}
sortVersionNumbers(vsns)
# [1] "1"        "1.1"      "1.1.1"    "1.1.1.1"  "1.1.1.2"  "1.1.1.10" "1.1.2"   
# [8] "1.1.10"   "1.2"      "1.10"     "10" 
sortVersionNumbers(sort(vsns))
# [1] "1"        "1.1"      "1.1.1"    "1.1.1.1"  "1.1.1.2"  "1.1.1.10" "1.1.2"   
# [8] "1.1.10"   "1.2"      "1.10"     "10" 

推荐答案

来自?numeric_version

> sort(numeric_version(vsns))
 [1] '1'        '1.1'      '1.1.1'    '1.1.1.1'  '1.1.1.2'  '1.1.1.10'
 [7] '1.1.2'    '1.1.10'   '1.2'      '1.10'     '10'  

看看这是如何实现的相对有趣. numeric_version将单个版本字符串拆分为整数部分,并将版本的向量存储为整数向量的列表. xtfrm上的一种方法(由sort()使用)将组成每个版本字符串的整数向量转换为数字值,胆量为

It's relatively interesting to see how this is implemented. numeric_version splits a single version string into integer parts, and stores the vector of versions as a list of integer vectors. A method on xtfrm (which is used by sort()) transforms the vector of integers making up each version string into a numeric value, with the guts being

base <- max(unlist(x), 0, na.rm = TRUE) + 1                                 
x <- vapply(x, function(t) sum(t/base^seq.int(0, length.out = length(t))), 
    1)

结果是一个数值向量,可用于以标准方式对原始向量进行排序.因此,临时解决方案是

the result is a numeric vector that can be used to order the original vector in a standard way. Thus an ad hoc solution is

xtfrm.my_version <- function(x) {
    x <- lapply(strsplit(x, ".", fixed=TRUE), as.integer)
    base <- max(unlist(x), 0, na.rm = TRUE) + 1
    vapply(x, function(t) sum(t/base^seq.int(0, length.out = length(t))), 1)
}

vsns  <- c("1", "10", "1.1", "1.10", "1.2", "1.1.1",
           "1.1.10", "1.1.2", "1.1.1.1", "1.1.1.10", "1.1.1.2")
class(vsns) = "my_version"
sort(vsns)

这篇关于排序包含语义版本的字符向量的文章就介绍到这了,希望我们推荐的答案对大家有所帮助,也希望大家多多支持IT屋!

查看全文
登录 关闭
扫码关注1秒登录
发送“验证码”获取 | 15天全站免登陆