【发布时间】:2017-07-13 09:12:30
【问题描述】:
我正在使用GA Package 来最小化一个函数。以下是我实施的几个阶段。
0。库和数据集
library(clusterSim) ## for index.DB()
library(GA) ## for ga()
data("data_ratio")
dataset2 <- data_ratio
set.seed(555)
1.二进制编码并生成初始种群。
initial_population <- function(object) {
## generate a population where for each individual, there will be number of 1's fixed between three to six
population <- t(replicate(object@popSize, {i <- sample(3:6, 1); sample(c(rep(1, i), rep(0, object@nBits - i)))}))
return(population)
}
2。适应度函数最小化 Davies-Bouldin (DB) 指数。
DBI2 <- function(x) {
## number of 1's will represent the initial selected centroids and hence the number of clusters
cl <- kmeans(dataset2, dataset2[x == 1, ])
dbi <- index.DB(dataset2, cl=cl$cluster, centrotypes = "centroids")
score <- -dbi$DB
return(score)
}
3.用户定义的交叉运算符。这种交叉方法将避免没有集群“打开”的情况。伪代码可以在here找到。
pairwise_crossover <- function(object, parents){
fitness <- object@fitness[parents]
parents <- object@population[parents, , drop = FALSE]
n <- ncol(parents)
children <- matrix(as.double(NA), nrow = 2, ncol = n)
fitnessChildren <- rep(NA, 2)
## finding the min no. of 1's between 2 parents
m <- min(sum(parents[1, ] == 1), sum(parents[2, ] == 1))
## generate a random int from range(1,m)
random_int <- sample(1:m, 1)
## randomly select 'random_int' gene positions with 1's in parent[1, ]
random_a <- sample(1:length(parents[1, ]), random_int)
## randomly select 'random_int' gene positions with 1's in parent[1, ]
random_b <- sample(1:length(parents[2, ]), random_int)
## union them
all <- sort(union(random_a, random_b))
## determine the union positions
temp_a <- parents[1, ][all]
temp_b <- parents[2, ][all]
## crossover
parents[1, ][all] <- temp_b
children[1, ] <- parents[1, ]
parents[2, ][all] <- temp_a
children[2, ] <- parents[2, ]
out <- list(children = children, fitness = fitnessChildren)
return(out)
}
4.突变。
k_min <- 2
k_max <- ceiling(sqrt(75))
my_mutation <- function(object, parent){
pop <- parent <- as.vector(object@population[parent, ])
for(i in 1:length(pop)){
if((sum(pop == 1) < k_max) && pop[i] == 0 | (sum(pop == 1) > k_min && pop[i] == 1)) {
pop[i] <- abs(pop[i] - 1)
return(pop)
}
}
}
5.把碎片放在一起。使用轮盘赌选择,交叉概率。 = 0.8,突变概率。 = 0.1
g2<- ga(type = "binary",
population = initial_population,
fitness = DBI2,
selection = ga_rwSelection,
crossover = pairwise_crossover,
mutation = my_mutation,
pcrossover = 0.8,
pmutation = 0.1,
popSize = 100,
nBits = nrow(dataset2))
我以这样一种方式创建了我的初始人口,即对于人口中的每个人,1's 的数量将固定在三到六之间。交叉和变异算子旨在确保解决方案最终不会有太多集群 (1's) 被“打开”。在集成之前,我已经分别尝试了我的交叉和突变功能,它们似乎工作正常。
理想情况下,最终的解决方案将具有来自初始群体的1's +-=1 的数量,即,如果一个个体的染色体中有三个1's,它最终会随机有两个、三个或四个1's。但我得到了这个解决方案,它显示 12 个集群 (1's) 正在“打开”,这意味着交叉和变异算子运行良好。
> sum(g2@solution==1)
[1] 12
这里的问题可以通过复制所有代码来重现。任何熟悉 GA 包的人都可以在这里帮助我吗?
[已编辑]
尝试使用不同的数据集iris,让我陷入了以下错误。 (仅更改数据,其余设置保持不变)
0。库和数据集
library(clusterSim) ## for index.DB()
library(GA) ## for ga()
## removed last column since it is a categorical data
dataset2 <- iris[-5]
set.seed(555)
> Error in kmeans(dataset2, centers = dataset2[x == 1, ]) :
initial centers are not distinct
我尝试查看code,发现此错误是由if(any(duplicated(centers))) 引起的。这可能意味着什么?
【问题讨论】:
-
你能分享你的样本初始人口数据吗?
-
它在
initial_population函数中,它将返回一个100 x 75 的二进制矩阵。我不太确定如何使用 ga 包访问它,但您可以使用population <- t(replicate(100, {i <- sample(3:6, 1); sample(c(rep(1, i), rep(0, 75 - i)))}))尝试一下 -
k_max和k_min? -
对不起,我错过了。
k_min = 2和k_max = ceiling(sqrt(75)) -
mutation_rate?还有其他未定义的常量吗?
标签: r genetic-algorithm mutation crossover