【发布时间】:2020-10-12 19:30:40
【问题描述】:
我有一个与我的 df 结构相关的技术问题。 它看起来像这样:
Month District Age Gender Education Disability Religion Occupation JobSeekers GMI
1 2020-01 Dan U17 Male None None Jewish Unprofessional workers 2 0
2 2020-01 Dan U17 Male None None Muslims Sales and costumer service 1 0
3 2020-01 Dan U17 Female None None Other Undefined 1 0
4 2020-01 Dan 18-24 Male None None Jewish Production and construction 1 0
5 2020-01 Dan 18-24 Male None None Jewish Academic degree 1 0
6 2020-01 Dan 18-24 Male None None Jewish Practical engineers and technicians 1 0
ACU NACU NewSeekers NewFiredSeekers
1 0 2 0 0
2 0 1 0 0
3 0 1 0 0
4 0 1 0 0
5 0 1 0 0
6 0 1 1 1
并且我正在寻找一种方法来在区域和求职者等 2 个变量之间进行卡方独立性检验,这样我就可以判断北部地区与求职者的关系是否比南部地区更多。 据我所知,数据结构有问题(地区是一个字符,求职者是一个整数,表示我有多少基于地区、性别、职业等的求职者) 我试图将其细分为像这样的地区和求职者:
Month District JobSeekers GMI ACU NACU NewSeekers NewFiredSeekers
<chr> <chr> <int> <int> <int> <int> <int> <int>
1 2020-01 Dan 33071 4694 9548 18829 6551 4682
2 2020-01 Jerusalem 21973 7665 3395 10913 3589 2260
3 2020-01 North 47589 22917 4318 20354 6154 3845
4 2020-01 Sharon 25403 6925 4633 13845 4131 2727
5 2020-01 South 37089 18874 2810 15405 4469 2342
6 2020-02 Dan 32660 4554 9615 18491 5529 3689
但它使处理变得更加困难 当然,我会接受任何其他可行的测试。
如果您需要更多信息,请帮助并告诉我,
莫舍
更新
# t test for district vs new seekers
# sorting
dist.newseek <- Cdata %>%
group_by(Month,District) %>%
summarise(NewSeekers=sum(NewSeekers))
# performing a t test on the mini table we created
t.test(NewSeekers ~ District,data=subset(dist.newseek,District %in% c("Dan","South")))
# results
Welch Two Sample t-test
data: NewSeekers by District
t = 0.68883, df = 4.1617, p-value = 0.5274
alternative hypothesis: true difference in means is not equal to 0
95 percent confidence interval:
-119952.3 200737.3
sample estimates:
mean in group Dan mean in group South
74608.25 34215.75
#wilcoxon test
# filtering Cdata to New seekers based on month and age
age.newseek <- Cdata %>%
group_by(Month,Age) %>%
summarise(NewSeekers=sum(NewSeekers))
#performing a wilcoxon test on the subset
wilcox.test(NewSeekers ~ Age,data=subset(age.newseek,Age %in% c("25-34","45-54")))
# Results
Wilcoxon rank sum exact test
data: NewSeekers by Age
W = 11, p-value = 0.4857
alternative hypothesis: true location shift is not equal to 0
方差分析测试
# Sorting occupation and month by new seekers
occu.newseek <- Cdata %>%
group_by(Month,Occupation) %>%
summarise(NewSeekers=sum(NewSeekers))
## Make the Occupation as a factor
occu.newseek$District <- as.factor(occu.newseek$Occupation)
## Get the occupation group means and standart deviations
group.mean.sd <- aggregate(
x = occu.newseek$NewSeekers, # Specify data column
by = list(occu.newseek$Occupation), # Specify group indicator
FUN = function(x) c('mean'=mean(x),'sd'= sd(x))
)
## Run one way ANOVA test
anova_one_way <- aov(NewSeekers~ Occupation, data = occu.newseek)
summary(anova_one_way)
## Run the Tukey Test to compare the groups
TukeyHSD(anova_one_way)
## Check the mean differences across the groups
library(ggplot2)
ggplot(occu.newseek, aes(x = Occupation, y = NewSeekers, fill = Occupation)) +
geom_boxplot() +
geom_jitter(shape = 15,
color = "steelblue",
position = position_jitter(0.21)) +
theme_classic()
【问题讨论】:
-
你不能做卡方。求职者是连续的。如果您想知道北部或南部是否与更多求职者相关,请进行 t.test 或 Manney-U ?
-
非常感谢您的回答@StupidWolf,您能解释一下如何使用当前表进行此特定的 t 测试吗?如果您发现很难做到,即使是具有相关示例的来源也会很棒。再次感谢您的关注!
-
我做到了。在主数据框中,每一行都基于年龄、宗教等多个部门,而不是求职者、新求职者等的总和,所以我试图减少部门,这样会更容易。整个数据框包含由冠状病毒引起的危机期间从 1 月到 4 月的详细信息,这是该项目的主要主题 - 失业率增加以及与职业、地区等的联系。希望这不是太多毫无价值的信息,再次感谢!
-
好吧..我认为你需要更多地考虑你的假设是什么。我试图提供一个关于你能做什么的快速答案。您还可以检查其他答案,它使用方差分析。理论上是一样的,但有更多的假设,是的,看看这对你的问题是否有意义
标签: r dataframe statistics chi-squared goodness-of-fit