文章检索
研究前沿

结合动态帧选择和时间运动增强的运动员动作识别

  • 刘芳
展开
  • 郑州经贸学院,郑州 451191
刘 芳(1978-),女,河南西平人,硕士,副教授,主要研究方向为动作识别、体育教育与训练。

收稿日期: 2025-09-26

  修回日期: 2025-11-07

  网络出版日期: 2026-07-14

基金资助

河南省科技攻关项目(252102210074)

Athlete Motion Recognition Combining Dynamic Frame Selection and Time Motion Enhancement

  • LIU Fang
Expand
  • Zhengzhou University of Economics and Business, Zhengzhou 451191, China

Received date: 2025-09-26

  Revised date: 2025-11-07

  Online published: 2026-07-14

摘要

为解决运动员动作识别领域中存在的骨骼数据标注成本高、帧选择不够合理和识别方法不够精确等缺陷,提出一种结合动态帧选择和时间运动增强的运动员动作识别方法。首先,基于视频RGB数据,提出一种结合全局和局部动态帧选择机制,并采用动态规划和自适应集束剪枝优化帧选择,确保全局一致性和局部多样性。然后,提出一种即插即用的时间和运动增强模块,可任意嵌入2D卷积神经网络,对视频动作特征进行学习。最后,在2个数据集上对所提识别方法的合理性和优越性进行验证,在HMDB51数据集上较次优性能提升3.16%,在PV数据集上较次优性能提升2.41%。实验表明:所提方法中结合全局和局部动态帧选择机制较为合理,能够兼顾全局一致性和局部多样性。同时即插即用的时间和运动增强模块能够有效学习视频中动作的时空关系,具有一定优越性。

本文引用格式

刘芳 . 结合动态帧选择和时间运动增强的运动员动作识别[J]. 复杂系统与复杂性科学, 2026 , 23(3) : 140 -151 . DOI: 10.13306/j.1672-3813.2026.03.017

Abstract

To address the shortcomings in the field of athlete movement recognition, such as high cost of bone data annotation, unreasonable frame selection, and insufficient accuracy of recognition methods, this paper proposes an athlete movement recognition method that combines dynamic frame selection and temporal motion enhancement. Firstly, based on video RGB data, a mechanism combining global and local dynamic frame selection is proposed, and dynamic programming and adaptive beam pruning are used to optimize frame selection to ensure global consistency and local diversity. Then, a plug-and-play time and motion enhancement module is proposed, which can be arbitrarily embedded into 2D convolutional neural networks to learn video action features. Finally, the rationality and superiority of the proposed recognition method are verified on two datasets. On the HMDB51 dataset, the performance is improved by 3.16% compared to the suboptimal performance, and on the PV dataset, it is improved by 2.41% compared to the suboptimal performance. The experiments show that the combination of the global and local dynamic frame selection mechanism in the proposed method is reasonable and can balance global consistency and local diversity. At the same time, the plug-and-play time and motion enhancement module can effectively learn the spatiotemporal relationship of actions in the video and has certain superiority.

参考文献

[1] GUO D, LI K, HU B, et al. Benchmarking micro-action recognition: dataset, methods, and applications[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(7): 6238-6252.
[2] REN B, LIU M, DING R, et al. A survey on 3d skeleton-based action recognition using learning method[J]. Cyborg and Bionic Systems, 2024, 5: 0100.
[3] SEDAGHATI N, ARDEBILI S, GHAFFARI A. Application of human activity/action recognition: a review[J]. Multimedia Tools and Applications, 2025: 1-30.
[4] 张宏宏, 李文华, 郑家毅, 等. 有人/无人机协同作战:概念、技术与挑战[J]. 航空学报, 2024, 45(15): 168-194.
ZHANG HH, LI W H, ZHENG J Y, et al. Manned/unmanned aerial vehicle cooperative combat system: concepts, technologies, and challenges[J]. Acta Aeronautica et Astronautica Sinica, 2024, 45(15): 168-194.
[5] KONG Y, FU Y. Human action recognition and prediction: a survey[J]. International Journal of Computer Vision, 2022, 130(5): 1366-1401.
[6] CHEN Z, MA F, LIU H, et al. MDT: A multiscale differencing transformer with sequence feature relationship mining for robust action recognition[J]. Applied Intelligence, 2025, 55(13): 1-19.
[7] 张祖习, 张战成, 胡伏原. 局部与长程时序互补建模的视频动作识别[DB/OL]. 计算机应用, 1-10[2025-08-02]. https://link.cnki.net/urlid/51.1307.TP.20250725.1050.004.
ZHANG Z X, ZHANG Z C, HU F Y. Local and long-range temporal complementary modeling for video action recognition[DB/OL]. Journal of Computer Applications, 1-10[2025-08-02]. https://link.cnki.net/urlid/51.1307.TP.20250725.1050.004.
[8] YAO D, CHEN H, YANG Z, et al. A dangerous behavior detection algorithm with the fusion of RGB data and skeleton information[J]. The Journal of Supercomputing, 2025, 81(13): 1290.
[9] CHEN S, LIU Y, ZHANG H, et al. A human location and action recognition method based on improved Yolov11 model[J]. Discover Artificial Intelligence, 2025, 5(1): 1-20.
[10] 张泽群, 李书婷, 曹其立, 等. 基于高阶时空自注意力的开集动作识别算法[DB/OL]. 计算机集成制造系统, 1-19[2025-08-20]. https://doi.org/10.13196/j.cims.2024.Z04.
ZHANG Z Q, LI S T, CAO Q L, et al. Open set action recognition algorithm based on high-order spatial-temporal self-attention[DB/OL]. Computer Integrated Manufacturing Systems, 1-19[2025-08-20]. https://doi.org/10.13196/j.cims.2024.Z04.
[11] RUAN X N, XIE B X, YIN Q X, et al. GLD: Global-local dynamic frame selection for action recognition[J]. IEEE Signal Processing Letters, 2025, 32: 2868-2872.
[12] HE K, ZHANG X, REN S, et al. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification[C]//Proceedings of the IEEE International Conference on Computer Vision. Santiago, Chile: IEEE, 2015: 1026-1034.
[13] 汪豪, 赵彬, 刘国华. 基于时间和运动增强的视频动作识别[J]. 吉林大学学报(工学版), 2025, 55(1): 339-346.
WANG H, ZHAO B, LIU G H. Temporal and motion enhancement for video action recognition[J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(1): 339-346.
[14] KUEHNE H, JHUANG H, GARROTE E, et al. HMDB: a large video database for human motion recognition[C]//2011 International Conference on Computer Vision. Barcelona, Spain: IEEE, 2011: 2556-2563.
[15] XIAN R Q, WANG X J, KOTHANDARAMAN D, et al. PMI Sampler: Patch similarity guided frame selection for aerial action recognition[C]//2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Honolulu, USA: IEEE, 2024: 6967-6976.
[16] WU W, WANG X, LUO H, et al. Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023: 6620-6630.
[17] LIU W, ZHONG X, JIA X, et al. Actor-aware alignment network for action recognition[J]. IEEE Signal Processing Letters, 2022, 29: 2597-2601.
[18] 张起尧, 桑海峰. 深度嵌套注意力下的SlowFast信息融合动作识别网络[J]. 电子测量与仪器学报, 2024, 38(3): 159-166.
ZHANG Q Y, SANG H F. SlowFast information fusion action recognition network based on deeply nested attention mechanism[J]. Journal of Electronic Measurement and Instrumentation, 2024, 38(3): 159-166.
[19] XING S, GUO Z, YU C, et al. Quality assessment of sports actions based on Adaptive-UniFormer[J]. Digital Signal Processing, 2025, 168: 105549.
[20] LIB, CHEN J, LI G, et al. Cross-modal contrastive masked autoencoder for compressed video pre-training[J]. IEEE Transactions on Image Processing, 2025, 34: 4500-4514.
[21] BERTASIUS G, WANG H, TORRESANI L. Is space-time attention all you need for video understanding?[C]//Proceedings of the 38th International Conference on Machine Learning. Virtual: PMLR, 2021: 813-824.
[22] TONG Z, SONG Y, WANG J, et al. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training[C]//Advances in Neural Information Processing Systems. New Orleans, USA: Curran Associates, Inc, 2022: 1-16.
文章导航

/

〈 〉