使用sciml包实现机器学习模型的交叉验证评估
在构建预测模型后,验证其在未见数据上的表现至关重要。当缺乏独立的外部验证数据集时,交叉验证提供了一种有效的评估方法。交叉验证(Cross-Validation,CV)是衡量模型泛化性能的常用技术,在数据科学和机器学习领域应用广泛。该方法将数据集划分为训练集和测试集,前者用于模型构建,后者用于评估模型性能。

k折交叉验证可以视为留一交叉验证的简化形式,它将原始数据随机划分为k个大小相近的子集(通常k取5-10)。每个子集依次作为测试集,其余k-1个子集合并为训练集,进行k次模型训练和评估。最终取各性能指标(如准确率、灵敏度、特异度、AUC等)的平均值作为模型性能的估计。这种方法通过平均多个评估结果,降低了数据划分方式对模型性能评估的影响。研究表明,当k取5或10时,模型评估的准确性和计算复杂性能达到较好的平衡。
10折交叉验证是将数据集均分为10个子集,每次使用9个子集作为训练数据,剩余1个子集作为测试数据,重复10次后计算平均性能指标,从而全面评估模型的泛化能力。

在sciml包中,我们开发了scikfoldcv函数,可以便捷地实现机器学习模型的10折交叉验证。下面通过一个实例演示其使用方法。
首先,加载必要的R包和数据:
library(scitable)
library(sciml)
library(httr)
library(base64enc)
# 读取不孕症数据集
infertility_data <- read.csv("E:/r/test/infertility.csv", sep=',', header=TRUE)

该数据集包含不孕症相关指标,其中8个变量中最后两个是PSM匹配结果,我们主要关注前6个变量:Education(教育程度)、age(年龄)、parity(产次)、induced(人流次数)、case(是否不孕,作为目标变量)、spontaneous(自然流产次数)。部分变量为分类变量,需要进行适当转换:
# 将分类变量转换为因子类型
infertility_data$education <- ifelse(infertility_data$education=="0-5yrs", 0,
ifelse(infertility_data$education=="6-11yrs", 1, 2))
infertility_data$spontaneous <- as.factor(infertility_data$spontaneous)
infertility_data$case <- as.factor(infertility_data$case)
infertility_data$induced <- as.factor(infertility_data$induced)
infertility_data$education <- as.factor(infertility_data$education)
假设我们要构建随机森林模型预测不孕症,并评估其AUC值。首先确定模型的输入变量:
predictors <- c("education", "age", "parity", "induced", "spontaneous")
使用scikfoldcv函数进行10折交叉验证仅需一行代码:
cv_results <- scikfoldcv(data=infertility_data, y="case", var=predictors,
type="randomForest", username=username, token=token)

从cv_results对象中提取测试数据和拟合模型:
test_data <- cv_results[["testlist"]]
model_list <- cv_results[["fitlist"]]
使用m.sciroc函数绘制ROC曲线,由于model_list是列表对象,需设置oblist = T参数:
library(scitable)
roc_output <- m.sciroc(model_list, newdata=test_data, oblist=T)
roc_output[["p"]]

由于验证集样本量有限,曲线不够平滑。我们可以提取各折的AUC值并在图中显示:
auc_values <- roc_output[["rocauc"]]
roc_output <- m.sciroc(model_list, newdata=test_data, legend.name=auc_values, oblist=T)
roc_output[["p"]]

至此,我们完成了可用于发表的随机森林模型10折交叉验证ROC曲线的绘制。该函数仍在持续优化中,更多功能请关注后续更新。
使用sciml包实现机器学习模型的交叉验证评估