法医学杂志 ›› 2026, Vol. 42 ›› Issue (2): 112-120.DOI: 10.12116/j.issn.1004-5619.2025.550701

• 论著 • 上一篇    下一篇

RNA测序数据归一化对法医学年龄推断的影响

赵晋远1,2(), 宣宇佳1,2, 陈安琪2, 路艳芳2, 廖梦筱2, 刘思彤2, 邢宇1,2, 王亚丽1, 陈丽琴1(), 李成涛1,2,3()   

  1. 1.内蒙古医科大学法医学教研室,内蒙古 呼和浩特 010110
    2.复旦大学法庭科学研究院,上海 200032
    3.司法鉴定科学研究院 上海市法医学重点实验室 司法部司法鉴定重点实验室 上海市司法鉴定专业技术服务平台,上海 200063
  • 收稿日期:2025-07-04 发布日期:2026-07-08 出版日期:2026-04-25
  • 通讯作者: 陈丽琴,李成涛
  • 作者简介:赵晋远(2000—),女,硕士研究生,主要从事法医遗传学研究;E-mail:1749916837@qq.com
  • 基金资助:
    国家自然科学基金重大项目(82293650);国家自然科学基金重大项目(82293654);国家自然科学基金青年科学基金资助项目(82302124)

Effect of RNA Sequencing Normalization on Forensic Age Estimation

Jinyuan ZHAO1,2(), Yujia XUAN1,2, Anqi CHEN2, Yanfang LU2, Mengxiao LIAO2, Sitong LIU2, Yu XING1,2, Yali WANG1, Liqin CHEN1(), Chengtao LI1,2,3()   

  1. 1.Department of Forensic Medicine, Inner Mongolia Medical University, Hohhot 010110, China
    2.Institute of Forensic Science, Fudan University, Shanghai 200032, China
    3.Shanghai Key Laboratory of Forensic Medicine, Key Laboratory of Forensic Science, Ministry of Justice, Shanghai Forensic Service Platform, Aca-demy of Forensic Science, Shanghai 200063, China
  • Received:2025-07-04 Online:2026-07-08 Published:2026-04-25
  • Contact: Liqin CHEN, Chengtao LI

摘要:

目的 评估4种常用RNA测序归一化方法在法医学年龄推断中的适用性,为优化年龄推断模型的准确性并提升其精确度提供参考依据。 方法 采集147例中国汉族无关个体的外周血样本,经RNA测序后采用每百万计数(counts per million,CPM)、每千个碱基的转录每百万映射读取的片段数(fragments per kilobase of transcript per million mapped reads,FPKM)、每百万转录本(transcripts per million,TPM)以及M值的修剪均值(trimmed mean of M-values,TMM) 4种方法对表达量进行归一化。采用Spearman相关性分析筛选与年龄相关的mRNA,并构建最小绝对收缩和选择算子(least absolute shrinkage and selection operator,LASSO)、支持向量机(support vector machine,SVM)、极端梯度提升(extreme gradient boosting,XGBoost) 3种年龄推断模型,比较4种归一化方法在建模中的表现。 结果 4种归一化方法筛选出的潜在年龄相关性标记数量不同,其中TMM归一化方法筛选出的标记数量最多,为2 912个。CPM和FPKM次之,分别筛选出2 897和1 481个潜在标记。TPM筛选出的标记数量最少,为338个。基于这些标记构建的年龄推断模型,训练集中平均绝对误差(mean absolute error,MAE)范围为4.38~8.62岁,测试集中MAE范围为5.85~9.29岁。TMM归一化方法在3种模型中表现最佳,尤其在XGBoost模型中,其训练集MAE为4.38岁,测试集MAE为5.85岁。 结论 不同归一化方法对法医学年龄推断模型的构建具有较大影响。TMM归一化方法能够有效提高预测精度,是法医学年龄推断的优选方法。

关键词: 法医遗传学, 年龄推断, RNA测序, 机器学习, 归一化

Abstract:

Objective To evaluate the applicability of four commonly used RNA sequencing normalization methods in forensic age estimation, and to provide a reference for optimizing the accuracy and precision of age estimation models. Methods Peripheral blood samples were collected from 147 unrelated Chinese Han individuals. After RNA sequencing, gene expression levels were normalized using four methods: counts per million (CPM), fragments per kilobase of transcript per million mapped reads (FPKM), transcripts per million (TPM), and trimmed mean of M-values (TMM). Spearman correlation analysis was used to screen age-related mRNAs. Three age estimation models were constructed using least absolute shrinkage and selection operator (LASSO), support vector machine (SVM), and extreme gradient boosting (XGBoost), respectively, to compare the performance of the four normalization methods. Results The four normalization methods screened different numbers of potential age-related markers. The TMM method screened the largest number of 2 912 markers. The number of potential markers screened was 2 897 and 1 481 for CPM and FPKM, respectively. The TPM method screened the fewest markers with the number of 338. Age estimation models constructed based on these markers exhibited mean absolute errors (MAE) ranging from 4.38 to 8.62 years in the training set and from 5.85 to 9.29 years in the test set. The TMM method performed the best among the three models, especially in the XGBoost model; it achieved an MAE of 4.38 years in the training set and 5.85 years in the test set for the XGBoost model. Conclusion Different normalization methods have a large impact on the construction of forensic age estimation models. The TMM normalization method can effectively improve prediction accuracy and is recommended as the preferred method for forensic age estimation.

Key words: forensic genetics, age estimation, RNA sequencing, machine learning, normalization

中图分类号: