Nurnoby, M Faisal (2026) Low-Latency Human Activity Recognition Agnostic Federated Framework Based on Multimodal Transformers. PhD thesis, King Fahd University of Petroleum and Minerals.
|
PDF
201706690_PhD_desertation_Faisal_ePrint.pdf Restricted to Repository staff only until 26 August 2027. Download (15MB) |
Arabic Abstract
يتطلب التعرّف على الأنشطة البشرية (HAR) في البيئات الذكية الواقعية نماذج تتسم بالدقة والكفاءة وانخفاض زمن الاستجابة والمحافظة على الخصوصية، وتكون قادرة على دمج المعلومات الواردة من مصادر بيانات متعددة. وتعالج هذه الأطروحة هذه المتطلبات من خلال تطوير إطار للتعلّم الاتحادي قائم على المحوّلات متعددة الوسائط، ويتميز بانخفاض زمن الاستجابة واستقلالية عن نوع النشاط، مع التركيز بصورة رئيسة على مراقبة السائق كأحد التطبيقات المتخصصة للتعرّف الآني على السلوك البشري في بيئات العمل الحرجة من حيث السلامة. وفي هذا التطبيق، يكون التركيز على كشف حالة عيني السائق وتثاؤبه ونعاسه بوصفها مهام ضمن مجال التعرّف على الأنشطة البشرية. يبدأ البحث بمراجعة منهجية وتحليل ببليومتري للدراسات السابقة القائمة على الرؤية الحاسوبية في مجال التعرّف على الأنشطة البشرية. وتُظهر هذه المراجعة تحولًا واضحًا من الشبكات العصبية الالتفافية التقليدية، والشبكات العصبية المتكررة، والأساليب القائمة على الرسوم البيانية، نحو الشبكات القائمة على المحوّلات، إلى جانب تزايد الاهتمام بالتعلّم متعدد الوسائط، والتعلّم الاتحادي المحافظ على الخصوصية. واستنادًا إلى هذه الملاحظات، تطوّر هذه الأطروحة نماذج مخصصة لمراقبة السائق، وتشمل نموذجًا تجميعيًا قائمًا على المحوّلات لكشف حالة العينين ومحوّلًا هجينًا زمانيًا مكانيًا لكشف التثاؤب. وبناءً على هذه المكونات، تقترح الأطروحة نموذجين متكاملين لكشف حالة النعاس للسائق، باعتبارها مهمة زمنية ذات مستوى أعلى ضمن مهام التعرّف على الأنشطة البشرية. ويتمثل النموذج الأول في محوّل رؤية هرمي متعدد المسارات، ومراعٍ للفروق بين الجنسين،، وقائم على النوافذ المزاحة لمعالجة منطقة الوجه. أما النموذج الثاني، فهو محوّل زماني مكاني مُجزّأ ذو آلية انتباه مفصولة تسلسليًا. ونظرًا إلى أن هذه النماذج، شأنها شأن معظم أساليب التعرّف على الأنشطة البشرية القائمة على الفيديو، تعالج بصورة مستمرة مقاطع زمنية ثابتة الطول، ومن ثم تستهلك موارد حاسوبية في معالجة إطارات زائدة، تقدم الأطروحة بعد ذلك منظومة كشف متسلسلة مُوجَّه بالأحداث، أي تقوم بالمعالجة الزمانية المكانية فقط عند رصد سلوك يحتمل أن يدل على النعاس. وتُدمج هذه المنظومة مؤشرات بصرية منخفضة التكلفة مستخرجة من إطارات مقتصة حول الوجه. كما تُنظَّم الانتقالات بين الحالات باستخدام آلية تباطؤ ذات عتبتين منفصلتين للتفعيل والإلغاء، بينما تتيح المعايرة الخاصة بكل فرد تخصيص عتبات المؤشرات وفقًا لخصائص السائق وظروف التسجيل. ولتعزيز الخصوصية وقابلية التوسع والتخصيص، يُوسّع الإطار ليشمل التعلّم الاتحادي باستخدام محوّلات زمنية خفيفة الوزن، تُدرَّب على متتاليات مدمجة لنسبتي أبعاد العين والفم، إلى جانب تمثيلات المخطط الطيفي. وأخيرًا، تُقدّم الأطروحة نموذجًا اتحاديًا هجينًا للرؤية واللغة، من خلال دمج حوار السائق بوصفه وسيطًا دلاليًا إضافيًا، بما يبيّن أن الدمج متعدد الوسائط يمكن أن يحسّن التعرّف على الأنشطة البشرية المراعي للسياق، مع الحفاظ على محلية البيانات. وقد أظهرت النتائج التجريبية المستمدة من مجموعتي بيانات معياريتين أن الأساليب المقترحة تتفوق باستمرار على نماذج مرجعية قوية؛ إذ حققت دقة بلغت 95.47 % في كشف نعاس السائق، إضافة إلى دقة بلغت 97.88 % في الكشف الاتحادي متعدد الوسائط لنعاس السائق باستخدام الرؤية واللغة.
English Abstract
Human Activity Recognition (HAR) in real-world smart environments requires models that are accurate, efficient, low-latency, privacy-preserving, and capable of combining information from multiple data sources. This dissertation addresses these require- ments by developing a low-latency event-driven activity-agnostic federated framework based on multimodal transformers, with a main focus on driver monitoring as a specialized HAR application. In this work, driver eye state detection, driver yawn detection, and driver drowsiness detection are treated as HAR tasks that support real-time behavioral understanding in safety-critical settings. The research starts with a systematic literature review and bibliometric analysis of vision-based HAR studies. This review shows a clear shift from conventional convolutional neural networks, recurrent neural networks, and graph-based methods toward transformer-based models, increasing interest in multimodal learning, and continuing research gaps in abnormal activity recognition, and privacy-preserving federated HAR. Based on these findings, the dissertation next develops HAR models for fine-grained driver monitoring, including a transformer ensemble for driver eye state detection and a hybrid spatiotemporal transformer for driver yawn detection. Building on these components, the dissertation then proposes two end-to-end models for driver drowsiness detection as a higher-level temporal HAR task: a gender-aware multi-stream shifted-window hierarchical vision transformer that captures localized facial regions while reducing demographic bias, and a factorized spatiotemporal transformer with sequentially disentangled attention that models both spatial appearance and temporal dynamics with low inference latency. Since these models, like most video-based HAR methods, process fixed-length temporal clips continuously and therefore spend computation on redundant frames, the dissertation next introduces an event-driven detection cascade that invokes spatiotemporal processing only when potentially drowsy behavior is observed. Inexpensive visual cues extracted from face-cropped frames are fused into a trigger confidence score, state transitions are governed by a hysteresis mechanism with separate trigger and release thresholds, an adaptive clip proposal module positions the temporal window according to the trigger confidence, and a spatiotemporal verification network confirms or rejects the proposed event; subject-wise calibration further personalizes the cue thresholds to individual drivers and recording conditions. To improve privacy, scalability, and personalization, the framework is further extended to federated learning through lightweight temporal transformers trained on compact eye and mouth aspect ratio sequences and spectrogram representations. Finally, a federated hybrid vision-language model is introduced by integrating driver dialogue as an additional semantic modality, showing that multimodal fusion can improve context-aware HAR while preserving data locality. Experimental results on two benchmark datasets reveal that the proposed methods consistently outperform strong baseline models, including 95.47% accuracy for driver drowsiness detection, 93.40% and 99.34% accuracy for driver yawn detection on YawDD and NTHU-DDD, 97.88% accuracy for federated multimodal vision-language drowsiness detection, and, for the event-driven cascade evaluated end to end on held-out NTHU-DDD subjects, 91.8% of drowsy clips detected at a 4.8% false-alarm rate with 97.8% precision and 93.5% balanced accuracy - confirming that restricting spatiotemporal computation to adaptively proposed windows preserves the temporal context required for reliable detection.
| Item Type: | Thesis (PhD) |
|---|---|
| Subjects: | Computer |
| Department: | College of Computing and Mathematics > Information and Computer Science |
| Thesis Advisor: |
Elsayed Elalfy,
|
| Thesis Committee Members: |
Tariq Ahmed Helmi Albusyuni,
Amir Hussain,
Abdul Jabbar Siddiqui,
Saeed Anwar,
|
| Depositing User: | M FAISAL NURNOBY |
| Date Deposited: | 27 Aug 2026 06:08 |
| Last Modified: | 27 Aug 2026 06:08 |
| URI: | https://eprints.kfupm.edu.sa/id/eprint/144746 |