The Road Ahead: A Comprehensive Review of Recent Advances in Traffic Sign and Lane Line Recognition for Autonomous Systems

الطريق إلى الأمام: مراجعة شاملة للتطورات الحديثة في التعرف على إشارات المرور ومسارات الحارات لأنظمة القيادة الذاتية

👤 Javier Santiago Olmos Medina, Jessica Gissella Maradey Lázaro, Anton Rassõlkin, Mahmoud Ibrahim 📄 IEEE Open Journal of Vehicular Technology, Vol. 7, 2025, pp. 160–178 🔗 10.1109/OJVT.2025.3635022 ✓ CC BY 4.0

الملخص

تُعد أنظمة الإدراك للتعرف على إشارات المرور (TSR) والتعرف على مسارات الحارات (LLR) ركيزتين أساسيتين للتشغيل الآمن والفعال لأنظمة مساعدة السائق المتقدمة (ADAS) والمركبات ذاتية القيادة بالكامل. تقدم هذه المراجعة تحليلاً شاملاً لأحدث الأبحاث الأكاديمية في هذه المجالات، مع التركيز الصارم على الأدبيات المنشورة من أكتوبر 2024 حتى الآن. يكشف التحليل عن العديد من الاتجاهات الرئيسية التي تشكل المجال. في TSR، يتميز التطور المعماري بتحسين الشبكات العصبية الالتفافية (CNNs)، وتخصص النماذج خفيفة الوزن القائمة على YOLO للتطبيقات المضمنة في الوقت الفعلي، وظهور الهندسات الهجينة CNN-Transformer. في الوقت نفسه، يُكرس جهد بحثي كبير لتعزيز المتانة ضد الظروف البيئية المعاكسة ومجموعة متزايدة من الهجمات الخصومة المتطورة والمقنعة فيزيائياً. في LLR، يتحول النموذج بسرعة من الكشف ثنائي الأبعاد في مستوى الصورة إلى التحديد المكاني ثلاثي الأبعاد الكامل والاستدلال الطوبولوجي، مدفوعاً بنماذج Transformer التي تتفوق في التقاط السياق العالمي والتبعيات بعيدة المدى. تشمل الموضوعات المشتركة بين كلا المجالين السعي الدؤوب لتحقيق الكفاءة الحسابية، والنهج المتمحور حول البيانات المتمثل في إنشاء معايير جديدة وصعبة للظروف المعاكسة والإدراك ثلاثي الأبعاد، والتكامل الناشئ ولكن التحويلي للتعلم متعدد المهام ونماذج الرؤية واللغة (VLMs) لبناء أنظمة قادرة على الاستدلال الشامل للمشهد. على الرغم من التقدم الكبير، لا تزال العديد من التحديات الرئيسية قائمة في مجال التعميم، لا سيما في معالجة الحالات النادرة طويلة الذيل وتطوير مقاييس تقييم مراعية للسلامة. من المتوقع أن تركز الأبحاث المستقبلية على التعلم ذاتي الإشراف، والتكامل الأقوى بين أنظمة الإدراك والتحكم، وتطوير الذكاء الاصطناعي الجدير بالثقة من خلال تحسين قابلية التفسير والمتانة. ستمهد هذه الجهود الطريق للجيل القادم من أنظمة المركبات الذكية.

1. المقدمة

يعتمد تطوير الأنظمة الذاتية، خاصة في قطاع السيارات، على قدرة المركبة على إدراك وتفسير بيئتها بدقة. من بين مهام الإدراك الأكثر أهمية التعرف على إشارات المرور (TSR) والتعرف على مسارات الحارات (LLR)، والتي تعمل كالعيون الرقمية للمركبة [1]. توفر هذه الأنظمة معلومات أساسية تدعم الوظائف عالية المستوى مثل التحديد الموضعي، وتخطيط المسار، والتحكم الطولي والعرضي في المركبة [2]. يضمن TSR الفعال الامتثال لقوانين وأنظمة المرور، مثل حدود السرعة وأوامر التوقف، بينما يمكّن LLR القوي من تحديد موضع المركبة بدقة داخل الطريق، مشكلاً الأساس لميزات مثل مساعد الحفاظ على المسار (LKA) والملاحة الذاتية. إن موثوقية هذه الأنظمة ليست مجرد مسألة أداء بل شرط أساسي لضمان سلامة ركاب المركبة والمشاة ومستخدمي الطريق الآخرين [3].

شهدت منهجيات TSR و LLR تحولاً عميقاً على مدى العقد الماضي، مدفوعاً إلى حد كبير بظهور التعلم العميق [4]. كانت الأساليب المبكرة متجذرة في تقنيات الرؤية الحاسوبية التقليدية، معتمدة على ميزات مصممة يدوياً وخطوط أنابيب معالجة متعددة المراحل. بالنسبة لـ TSR، غالباً ما تضمن ذلك تجزئة قائمة على الألوان (مثل تحديد المناطق الحمراء أو الزرقاء) تليها تحليل الشكل وتصنيف باستخدام واصفات الميزات مثل الرسم البياني للتدرجات الموجهة (HOG) أو الأنماط الثنائية المحلية (LBP) [3]. وبالمثل، استخدمت خوارزميات LLR التقليدية تقنيات مثل كشف الحواف Canny، وتصفية الفضاء اللوني، وتحويل Hough لتحديد الخطوط الملائمة لعلامات الحارات في الصورة. على الرغم من أن هذه الأساليب كانت تأسيسية، إلا أنها غالباً ما واجهت صعوبات مع التنوع الهائل لظروف القيادة في العالم الحقيقي، بما في ذلك تغير الإضاءة، والطقس العكسي، والانسدادات، وتدهور الإشارات [1].

لقد حل ظهور التعلم العميق (DL) إلى حد كبير محل هذه الأساليب التقليدية، حيث يقدم حلولاً شاملة مبنية على البيانات بأداء وقدرات تعميم فائقة بشكل كبير [4]. يمكن وصف الإطار الزمني الأخير الذي تركز عليه هذه المراجعة ليس كفترة اكتشاف أولي بل كفترة نضوج وتخصص ومواجهة نقدية للتحديات المستمرة في النشر في العالم الحقيقي. لقد تجاوز المجتمع البحثي مجرد إظهار دقة عالية على مجموعات البيانات القياسية النظيفة لمعالجة المشكلات الأكثر دقة وصعوبة التي لا تزال قائمة [5]، [6]، [7].

الهدف البحثي لهذه المراجعة هو تجميع وتحليل أحدث ما توصلت إليه التطورات في التعرف على إشارات المرور (TSR) والتعرف على مسارات الحارات (LLR) بشكل منهجي، مع التركيز على التحولات المعمارية والمنهجية والمفاهيمية التي حددت المجال في الفترة الأخيرة. لتنظيم هذا التحليل، تتناول الورقة ثلاثة أسئلة بحثية مركزية: (1) ما هي الابتكارات المعمارية الأساسية، وتقنيات تعزيز المتانة، والمخاوف الأمنية الخصومة التي تشكل تطور أنظمة TSR الحديثة؟ (2) كيف يتحول نموذج التعرف على مسارات الحارات من الكشف ثنائي الأبعاد نحو الفهم المكاني ثلاثي الأبعاد والطوبولوجي، وما هي التقنيات التمكينية الرئيسية الدافعة لهذا التحول؟ (3) ما هي الاتجاهات التكنولوجية الشاملة المشتركة بين كل من TSR و LLR، وكيف تؤثر على مسار مجال إدراك المركبات بأكمله؟ تجيب الورقة على هذه الأسئلة من خلال الأقسام التالية، التي تتناول التقدم في TSR، والنماذج المتطورة في LLR، والاتجاهات الناشئة، وأخيراً التوليف والتحديات المفتوحة.

2. التقدم في التعرف على إشارات المرور (TSR)

يتطور البحث الحديث في TSR على طول تسلسل هرمي واضح للتجريد. المستوى الأساسي يعالج المتانة البصرية—تحدي "رؤية الإشارة بوضوح" من خلال التغلب على التدهور الناجم عن ظروف الليل أو الضباب أو التشويش. المستوى التالي يعالج التعميم الدلالي والجغرافي، بهدف فهم "ما تعنيه الإشارة" بغض النظر عن اختلافات التصميم عبر البلدان المختلفة. يركز المستوى الأعلى على الاستدلال السياقي، الذي لا يتطلب فقط التعرف على الإشارة بل تحديد "كيف تنطبق قاعدتها على الطريق وإجراءات المركبة"، وهي مشكلة تكامل متعدد الوسائط. يتوج هذا التسلسل الهرمي ضرورة الثقة والأمن، والتي تعالج المتانة الخصومة وقابلية التفسير (XAI) لبناء أنظمة قابلة للتحقق والتوثيق في النهاية.

أ. الابتكارات المعمارية: السعي وراء السرعة والدقة

ركز البحث الحديث في TSR على النماذج الشاملة التي تتعلم خط الأنابيب بأكمله من البكسلات الخام إلى تحديد موقع الإشارة وتصنيفها، متجاوزة الأساليب التقليدية متعددة المراحل. جوهر أي نظام TSR حديث هو هندسة شبكته العصبية. يُظهر البحث الحديث نهجاً متعدد الجوانب حيث يتم اتباع استراتيجيات مختلفة لتحسين المفاضلات بين الدقة وسرعة الاستدلال وحجم النموذج. يمكن تصنيف هذا التطور بشكل عام إلى هندسات خفيفة الوزن للنشر في الوقت الفعلي، ونماذج عالية الدقة لأقصى أداء، وتكامل قابلية التفسير كمبدأ تصميم أساسي.

يركز جزء كبير من البحث على الهندسات الشاملة المحسّنة للنشر على الأنظمة المضمنة محدودة الموارد الموجودة في المركبات [8]. بالنسبة للتطبيقات التي تتطلب أداءً في الوقت الفعلي، تظل كاشفات المرور الواحدة مثل عائلة YOLO خياراً مهيمناً بسبب تصميمها أحادي المرور الذي يعطي الأولوية للسرعة [9]. ركز البحث الحديث على تكييف وتعزيز متغيرات YOLO لمواجهة التحديات المحددة لـ TSR. على سبيل المثال، يعيد نموذج AF-YOLO بناء العمود الفقري لـ YOLOv5 باستخدام وحدة C3-SE وشبكة هرم الميزات المقاربة (AFPN) [10]. يقلل هذا التصميم بشكل كبير من عدد معلمات النموذج بنسبة 21.79٪ والتعقيد الحسابي بنسبة 3.73٪ مقارنة بـ YOLOv5s الأساسي، مع زيادة متوسط الدقة (mAP) بنسبة 1٪ وتعزيز سرعة الكشف بمقدار 16 إطاراً في الثانية (FPS) [10].

بالتوازي مع السعي لتحقيق الكفاءة، يستمر خط آخر من البحث في دفع حدود الأداء. تظل CNNs الأساس المتين لأنظمة TSR عالية الأداء. تستمر الدراسات في اقتراح هندسات CNN عميقة مصممة بدقة تحقق دقة شبه كاملة، تتجاوز في بعض الحالات 99.63٪ على GTSRB و 99.68٪ على CTSD [2]. النماذج القائمة على الهندسات القوية مثل EfficientDet، والتي تستخدم دمج الميزات المتقدم مثل BiFPN، أظهرت أيضاً أداءً عالياً، محققة mAP@0.5 بنسبة 96.6٪ على TT100K [15]. تعكس الاتجاه الأوسع في الرؤية الحاسوبية، يتم استكشاف هندسات Transformer بشكل متزايد. ومع ذلك، فإن تكلفتها الحسابية العالية حولت التركيز نحو النماذج الهجينة التي تجمع بين نقاط قوة CNNs و Transformers. مثال رئيسي هو E-MobileViT، نموذج خفيف الوزن لـ TSR يبني على MobileViT، وهو هجين CNN-ViT موجود [16].

ب. تحليل مقارن لابتكارات TSR المعمارية

يقدم الجدول الأول مقارنة لابتكارات TSR المعمارية الحديثة. تُظهر التطورات في التعرف على إشارات المرور مكاسب ملحوظة في الدقة والكفاءة والقدرة على التكيف. تحقق نماذج مثل Real-time TSR CNN و E-MobileViT دقة شبه كاملة (حوالي 99.9٪) على المعايير مع البقاء محسّنة للنشر في الوقت الفعلي أو على الأجهزة المحمولة. يقدم AF-YOLO فوزاً ثلاثياً نادراً—mAP أعلى، واستدلال أسرع (345 FPS)، ومعلمات أقل—مما يجعله مثالياً للأنظمة المضمنة. يحقق TSD-Net أحدث النتائج على TT100K لتحدي الأشياء الصغيرة، بينما يمزج YOLO-RCNN بين سرعة YOLO وتنقية R-CNN للكشف عن الإشارات الحرجة. من المهم أن نموذج Explainable CNN يضيف الشفافية عبر Grad-CAM، مما يسد الفجوة بين الدقة وقابلية التفسير [1]، [10]، [12]، [16]، [23].

ج. تعزيز المتانة والتعميم

على الرغم من أن تحقيق دقة عالية على مجموعات البيانات القياسية يعد معلماً مهماً، إلا أن الاختبار الحقيقي لنظام TSR هو قدرته على الأداء بشكل موثوق في الظروف غير المتوقعة وغير المثالية للقيادة في العالم الحقيقي. يُكرس جزء كبير من الأبحاث الحديثة لتعزيز متانة النموذج ضد الظروف البيئية المعاكسة والسيناريوهات البصرية الصعبة.

تعد العوامل البيئية مثل ضعف الإضاءة والطقس العكسي (المطر والضباب) والعراقيل البصرية من المساهمين الرئيسيين في فشل أنظمة الإدراك. يعالج الباحثون هذه المشكلات من زوايا متعددة. أحد الاستراتيجيات الأساسية هو زيادة البيانات المتطورة—محاكاة مجموعة واسعة من التدهورات في العالم الحقيقي مثل الضوضاء والانسدادات وظروف الطقس المختلفة [20]. بالإضافة إلى ذلك، للتعامل مع التحديات التي لا يمكن حلها بواسطة مستشعر واحد، تقدم الأنظمة متعددة الوسائط حالاً واعداً، كما هو موضح في دراسة تستخدم منصة NVIDIA Jetson Nano مع إعداد كاميرا مزدوجة (كاميرا CSI نهارية وأشعة تحت حمراء ليلية) [20]. نهج آخر هو استخدام خوارزميات إعادة بناء فائقة الدقة لتعزيز الإشارات الباهتة أو المشوشة قبل التعرف عليها [21].

يبقى اكتشاف الإشارات الصغيرة (بسبب المسافة) أو المغطاة جزئياً تحدياً صعباً. صمم الباحثون هندسات شبكات متخصصة مثل TSD-Net، الذي يستخدم الالتفاف الديناميكي وفرع كشف عالي الدقة محسّن [23]. كما أن تحسين دمج الميزات عبر المقاييس المختلفة من خلال رؤوس الكشف مثل ADFF أمر بالغ الأهمية للكشف المتزامن عن الإشارات الكبيرة القريبة والصغيرة البعيدة [10]، [23].

د. جبهة الأمن: الهجمات الخصومة والدفاعات

مع ازدياد قدرة أنظمة TSR وانتشارها، تصبح أهدافاً أكثر جاذبية للمهاجمين الخبيثين. يشكل ضعف الشبكات العصبية العميقة أمام الهجمات الخصومة خطراً كبيراً على السلامة والأمن [5]. يشهد البحث الحديث سباق تسلح بين تطوير أساليب هجوم أكثر تطوراً وقابلية للتنفيذ فيزيائياً وإنشاء آليات دفاع أكثر قوة.

على صعيد الهجوم، يبحث الباحثون في الاضطرابات "الطبيعية" التي تحاكي الظواهر البيئية الشائعة مثل الظلال وبقع الضوء [35]. تشكل الملصقات الفيزيائية تهديداً خطيراً، حيث أظهرت الأبحاث هجوماً "شاملاً" حيث يمكن لملصق واحد أو اثنين أن يسببا تصنيفاً خاطئاً لأي نوع من الإشارات [36]. يأخذ هجوم FIGhost التخفي إلى مستوى أعلى باستخدام حبر فلوري غير مرئي تحت الإضاءة العادية ويتم تنشيطه فقط بالأشعة فوق البنفسجية [37].

على صعيد الدفاع، تشمل الاستراتيجيات تعديلات معمارية (مثل شبكات التحويل المكاني [35])، وتحويل بيانات الإدخال (مثل التدريب متعدد الدقة الذي يقوم بالتصفية المنخفضة للضوضاء الخصومة [38])، والدفاعات القائمة على إعادة البناء (مثل autoencoders ثنائية الوضع التي تكتشف الهجمات وتصلحها [39])، ومنهجيات التدريب القوية (مثل استخدام LoRA لتصحيح نماذج ViT المدربة مسبقاً بكفاءة [40]).

3. النماذج المتطورة في التعرف على مسارات الحارات (LLR)

شهد LLR، حجر الزاوية في ADAS الحديث والقيادة الذاتية، فترة من التطور المنهجي السريع. يتحرك المجال بشكل حاسم إلى ما وراء تحليل الصور ثنائي الأبعاد التقليدي نحو نماذج أكثر تطوراً تحتضن التعلم الشامل والوعي المكاني ثلاثي الأبعاد والدمج متعدد الوسائط للمستشعرات. هذه التطورات مدفوعة بالحاجة إلى دقة ومتانة أكبر في مواجهة هندسات الطرق المعقدة والظروف البيئية الصعبة.

أ. مراجعة منهجية لكشف الحارات ثنائي الأبعاد

تتضمن مهمة كشف الحارات ثنائي الأبعاد تحديد حدود ممرات القيادة داخل صورة ثنائية الأبعاد تلتقطها كاميرا المركبة. بينما بقي الهدف ثابتاً، تحولت أساليب تحقيقه بشكل كبير. كانت خطوط أنابيب LLR التقليدية عادةً عمليات متعددة المراحل قائمة على الاستدلال—تحويل التدرج الرمادي، والتشويش Gaussian، وكشف الحواف Canny، وتحويل Hough. بينما هذه الطرق بسيطة حسابياً، أداؤها هش ويعتمد بشكل كبير على علامات حارات واضحة وظروف إضاءة مستقرة.

النهج الحديث، الذي يهيمن عليه التعلم العميق، انتقل نحو النماذج الشاملة التي تتعلم كشف الحارات مباشرة من بيانات البكسل الخام، مما يقلل الاعتماد على المعلمات المعدلة يدوياً والخطوات الوسيطة الهشة [53]. يمكن تصنيف هذه النماذج الشاملة إلى عدة نماذج متنافسة: الأساليب القائمة على التجزئة (تصنيف على مستوى البكسل)، والأساليب القائمة على المراسي (خطوط مرساة محددة مسبقاً)، والأساليب البارامترية (ملاءمة متعددة الحدود أو منحنيات Bézier)، والأساليب القائمة على Transformer.

التطور الأكثر أهمية في LLR ثنائي الأبعاد هو تطبيق هندسات Transformer. تسمح آلية الانتباه الذاتي في صميم Transformers بالتقاط السياق العالمي ونمذجة التبعيات بعيدة المدى عبر الصورة بأكملها—ميزة حاسمة لكشف الحارات تمكن النموذج من استنتاج مسار الحارة حتى عندما تكون مغطاة جزئياً [4]. تشمل النماذج المبتكرة القائمة على Transformer MHFS-FORMER، الذي يقدم ميزات هجينة متعددة المقاييس محسّنة ووحدة انتباه قابلة للتشوه متعددة المرجع مصممة لالتقاط البنية الممدودة للحارات [4]، و CCHA-Net، الذي يقترح آلية انتباه هجينة عبر الالتفاف تحقق درجة F1 عالية جداً تبلغ 80.2٪ على مجموعة بيانات CULane الصعبة [56].

ب. تجاوز الحدود: كشف الحارات ثلاثي الأبعاد ومتعدد الوسائط

بينما LLR ثنائي الأبعاد مجال ناضج، فإن محدوديته الأساسية هي أنه يوفر فقط موضع الحارات في مستوى الصورة. للوظائف الذاتية المتقدمة، تحتاج المركبة إلى معرفة الموقع المكاني ثلاثي الأبعاد الدقيق للحارات في العالم الحقيقي [6]. هذا دفع كشف الحارات ثلاثي الأبعاد إلى طليعة أبحاث LLR [67].

يتضمن النهج الشائع تحويل الميزات من الصورة الأمامية (FV) إلى تمثيل منظور عين الطائر (BEV). ومع ذلك، فإن هذا التحويل يمكن أن يكون مكلفاً حسابياً وعرضة للأخطاء، خاصة على أسطح الطرق غير المستوية. لتجاوز هذه المشكلات، ظهر اتجاه قوي لتطوير طرق خالية من BEV تتنبأ بالحارات ثلاثية الأبعاد مباشرة من ميزات FV. نموذج Anchor3DLane++ هو طريقة حديثة خالية من BEV تحدد مراسي الحارات مباشرة في الفضاء ثلاثي الأبعاد [67]. ابتكاره الرئيسي هو وحدة توليد المراسي التكيفية القائمة على النموذج الأولي (PAAG)، التي تولد ديناميكياً مجموعة متفرقة من المراسي المتوافقة جيداً مع الحارات في صورة الإدخال.

يمثل الحارة ثلاثية الأبعاد قرار تصميم حاسم.許多 طريقة، سعياً للكفاءة الحسابية، تصمم الحارات كمجموعة متفرقة من النقاط ثلاثية الأبعاد. ومع ذلك، كشف التحليل النقدي الحديث أن هذا التمثيل يمكن أن يكون معيباً—عملية توليد التسميات الأساسية غالباً ما تقطع الحارات عند نقاط نهايتها المرئية، مما قد يؤدي إلى أخطاء تصل إلى 20 متراً [6]. لمعالجة هذا، تقترح طرق جديدة استراتيجية الترقيع ورأس EP (نقطة النهاية) الذي يتنبأ صراحة بالمسافة المقطوعة لتمثيل بنية الحارة الكاملة بدقة أكبر. بالإضافة إلى ذلك، حفزت التحديات المتأصلة للأنظمة القائمة على الكاميرات فقط في تقدير العمق البحث في الدمج متعدد الوسائط، والجمع بين الكاميرات و LiDAR، كما هو موضح في نموذج FusionLane [68].

ج. ما وراء العلامات: الحارات الافتراضية والملاحة بدون حارات

يقودنا البحث في الطرق غير المحددة إلى التحدي الأكثر تعقيداً: الملاحة في بيئات المرور الخالية من الحارات، مثل الساحات المفتوحة أو التقاطعات الكبيرة [71]. في هذه السيناريوهات، يتحول النموذج بالكامل من "كشف الحارات" إلى التوليد الديناميكي والملاحة عبر "الحارات الافتراضية" [72]. يتطلب هذا من المركبة ليس فقط إدراك الطريق، بل استنتاج مسار آمن وفعال بناءً على سلوك الوكلاء الآخرين وسياق الطريق الأساسي. يجري تطوير نماذج سلوكية جديدة مثل "المتابعة بدون حارات" (LFF) [73]، وتتم معالجة اتخاذ القرار بشكل متزايد بواسطة أطر التعلم المعزز متعدد الوكلاء (MARL) [74]، [75].

4. الاتجاهات الناشئة في هندسات الإدراك

بينما TSR و LLR مهمتان متميزتان، فإن تقدمهما مدفوع بمحرك دوري مشترك وقوي ذاتي التعزيز يشكل المجال بأكمله. تبدأ هذه الدورة عندما تمكن الأجهزة المتخصصة الأكثر قوة (SoCs, FPGAs) من إنشاء خوارزميات أكثر تعقيداً، مثل Transformers. هذه الخوارزميات بدورها "تحل" أو تشبع المعايير الحالية، مما يجبر المجتمع على إنشاء مجموعات بيانات جديدة وأصعب. التحديات الجديدة ثم تطلب خوارزميات أكثر تطوراً وأجهزة أكثر قوة، مكملة الدورة. نتيجة حاسمة لهذه الحلقة هي خلق "فجوة حسابية" متسعة بين الأبحاث المتطورة وواقع النشر في السوق الشامل المحدود التكلفة والطاقة.

أ. السعي لتحقيق الكفاءة: نماذج خفيفة للأنظمة المضمنة

الشرط العالمي لكلا نظامي TSR و LLR هو القدرة على العمل في الوقت الفعلي على منصات الحوسبة محدودة الموارد الموجودة في المركبات [16]. أدى ذلك إلى تركيز مشترك على تطوير نماذج خفيفة الوزن وفعالة لا تضحي بالدقة. يتم معالجة هذا السعي للكفاءة من خلال عدة استراتيجيات متقاربة: استخدام الالتفافات القابلة للفصل حسب العمق، والاستخدام الاستراتيجي لآليات الانتباه [10]، [16]، والشبكات العصبية الثنائية (BNNs)، والتعلم متعدد المهام (MTL). نموذج YOLOP، على سبيل المثال، يؤدي كشف الأشياء المرورية وتجزئة المنطقة القابلة للقيادة وكشف الحارات ضمن إطار موحد واحد [79].

ب. الثورة المتمحورة حول البيانات: مجموعات البيانات والمعايير

يرتبط وتيرة التقدم في كل من TSR و LLR ارتباطاً وثيقاً بتوفر مجموعات بيانات عالية الجودة وواسعة النطاق ومتنوعة. الاتجاه الواضح في الأدبيات الحديثة هو نهج متمحور حول البيانات للبحث، حيث يُنظر إلى إنشاء معايير جديدة وأكثر تحدياً كمحرك رئيسي للابتكار. تركز هذه المجموعات الجديدة على التقاط البيانات من الظروف المعاكسة (FoggyLane، FoggyCULane [84])، والتحديات الليلية (ZND Dataset [82])، والإدراك ثلاثي الأبعاد (OpenLane-V2 [7])، والتنوع الإقليمي (مجموعة بيانات الطرق الهندية [85]، مجموعة بيانات بنغلاديش [86])، ومجموعات البيانات الاصطناعية [83].

ج. صعود الإدراك المتكامل القائم على الاستدلال

الاتجاه الأكثر استشرافاً هو التحرك نحو أنظمة إدراك أكثر تكاملاً وقادرة على الاستدلال، بدلاً من مجرد أداء مهام التعرف على الأنماط المعزولة. يمثل التعلم متعدد المهام (MTL) الخطوة الأولى نحو هذا التكامل. تطور أعمق هو تطبيق نماذج الرؤية واللغة (VLMs) لتعزيز إدراك المركبة. هذا ينتقل إلى ما وراء التصنيف البسيط نحو الاستدلال الحقيقي. على سبيل المثال، يستخدم نظام LKAlert VLM ليس فقط لكشف الحارات ولكن أيضاً للتنبؤ بالفشل الوشيك لنظام مساعد الحفاظ على المسار وتوليد تفسيرات باللغة الطبيعية للسائق [94]. يمثل تقديم قدرات الاستدلال تحولاً نموذجياً محتملاً من الإدراك كتعرف على الأنماط إلى الإدراك كاستدلال.

5. التوليف والتحديات المفتوحة والتوجهات المستقبلية

كانت الفترة الأخيرة من البحث في التعرف على إشارات المرور ومسارات الحارات، الممتدة من أواخر 2024 إلى 2025، فترة من التقدم الديناميكي والتصاعد المتزايد. المجال ينضج بعد ثورة التعلم العميق الأولية ويكافح الآن المشكلات الدقيقة والصعبة التي تقف في طريق النشر الواسع والآمن والموثوق. يكشف الأدب عن سعي مزدوج: من ناحية، يدفع الباحثون سقف الأداء بنماذج أكثر تعقيداً وقدرة؛ ومن ناحية أخرى، يعملون على خفض عتبة النشر بتطوير هندسات عالية الكفاءة وخفيفة الوزن.

على الرغم من التقدم الملحوظ، لا تزال العديد من التحديات الكبيرة قائمة. يعد تعميم المجال (Domain Generalization) مشكلة حرجة—فالنماذج المدربة على مجموعة بيانات واحدة غالباً ما تشهد انخفاضاً كبيراً في الأداء عند نشرها في منطقة مختلفة. مشكلة الذيل الطويل (Long Tail) المتعلقة بالحالات النادرة غير المتوقعة لا تزال تشكل تحدياً. هناك حاجة ملحة لتطوير مقاييس تقييم مراعية للسلامة تعكس بشكل أفضل المخاطر في العالم الحقيقي. كما أن ندرة بيانات العالم الحقيقي وتكلفة التصنيف تشكلان عنق زجاجة كبير، خاصة للمهام المتقدمة مثل الحارات ثلاثية الأبعاد.

تشير التحديات المحددة إلى عدة اتجاهات بحثية مستقبلية واعدة: التعلم ذاتي الإشراف ومدى الحياة للتغلب على ندرة البيانات؛ التكامل الأوثق بين الإدراك والتحكم؛ النماذج التوليدية ونماذج العالم لإنشاء بيانات تركيبية واقعية؛ وتطوير الذكاء الاصطناعي الجدير بالثقة من خلال المتانة والأمن وقابلية التفسير [1].

6. الاستنتاجات

قامت هذه المراجعة بتحليل منهجي لأحدث ما توصلت إليه التطورات في التعرف على إشارات المرور ومسارات الحارات في الإطار الزمني الأخير من خلال معالجة ثلاثة أسئلة بحثية مركزية. فيما يتعلق بتطور أنظمة TSR، وجدت المراجعة أن التقدم يُحدد بثلاثة مجالات رئيسية: معماريًا، يظهر المجال نهجاً متعدد الجوانب مع التحسين المستمر لشبكات CNN عالية الدقة وتخصص نماذج YOLO خفيفة الوزن وظهور نماذج هجينة CNN-Transformer؛ في المتانة، تركز الجهود على زيادة البيانات المتطورة والحلول متعددة الوسائط؛ وفي جبهة الأمن، يتميز المجال بـ "سباق تسلح" بين الهجمات المادية المتزايدة التخفي والدفاعات الأكثر كفاءة وقدرة على التكيف.

فيما يتعلق بتحول نموذج LLR، يؤكد التحليل تحركاً حاسماً بعيداً عن الكشف ثنائي الأبعاد نحو التحديد المكاني ثلاثي الأبعاد والاستدلال الطوبولوجي، مدعوماً بهندسات Transformer والطرق الخالية من BEV والدمج متعدد الوسائط. أخيراً، فيما يتعلق بالاتجاهات الشاملة، حددت المراجعة ثلاثة موضوعات رئيسية تشكل كلاً من TSR و LLR: السعي للكفاءة الحسابية، والنهج المتمحور حول البيانات، والتحول نحو الإدراك المتكامل القائم على الاستدلال عبر VLMs. بينما المهام الأساسية محلولة جيداً تحت الظروف القياسية، تركز جبهة البحث الحالية على تحقيق المتانة والأمن والفهم الحقيقي ثلاثي الأبعاد للمشهد، مما يمهد الطريق للجيل القادم من أنظمة المركبات الذكية الجديرة بالثقة.

المراجع

قائمة المراجع الكاملة (أكثر من 94 مصدراً) متاحة في الملف الأصلي بصيغة PDF على IEEE Xplore. المقال منشور بموجب ترخيص CC BY 4.0.

Abstract

The perception systems for Traffic Sign Recognition (TSR) and Lane Line Recognition (LLR) are foundational pillars for the safe and effective operation of Advanced Driver-Assistance Systems (ADAS) and fully autonomous vehicles. This review provides a comprehensive analysis of the latest academic research in these domains, strictly focusing on literature published from October 2024 to the present. The analysis reveals several key trends shaping the field. In TSR, architectural evolution is characterized by the refinement of Convolutional Neural Networks (CNNs), the specialization of light-weight YOLO-based models for real-time embedded applications, and the emergence of hybrid CNN-Transformer architectures. Concurrently, a significant research thrust is dedicated to enhancing robustness against environmental adversities and a growing spectrum of sophisticated, physically plausible adversarial attacks. In LLR, the paradigm is rapidly shifting from 2D image-plane detection to full 3D spatial localization and topology reasoning, driven by Transformer-based models that excel at capturing global context and long-range dependencies. Cross-cutting themes common to both domains include a relentless drive for computational efficiency, a data-centric approach marked by the creation of new, challenging benchmarks for adverse conditions and 3D perception, and the nascent but transformative integration of multi-task learning and Vision-Language Models (VLMs) to build systems capable of holistic scene reasoning. Despite significant progress, several key challenges persist in the field of domain generalization, particularly in handling long-tail corner cases and developing safety-aware evaluation metrics. Future research is expected to focus on self-supervised learning, stronger integration between perception and control systems, and the advancement of trustworthy AI through improved explainability and robustness. These efforts will lay the groundwork for the next generation of intelligent vehicle systems.

1. Introduction

The development of autonomous systems, especially in the automotive sector, hinges on the ability of a vehicle to accurately perceive and interpret its environment. Among the most critical perception tasks are Traffic Sign Recognition (TSR) and Lane Line Recognition (LLR), which serve as the digital eyes of the vehicle [1]. These systems provide fundamental information that underpins higher-level functions such as localization, path planning, and longitudinal and lateral vehicle control [2]. Effective TSR ensures compliance with traffic laws and regulations, such as speed limits and stop commands, while robust LLR enables precise vehicle positioning within the roadway, forming the basis for features like Lane Keeping Assist (LKA) and autonomous navigation. The reliability of these systems is not merely a matter of performance but a prerequisite for ensuring the safety of vehicle occupants, pedestrians, and other road users [3].

The methodologies for TSR and LLR have undergone a profound transformation over the past decade, largely driven by the advent of deep learning [4]. Early approaches were rooted in classical computer vision techniques, relying on hand-crafted features and multi-stage processing pipelines. For TSR, this often involved color-based segmentation (e.g., identifying red or blue regions) followed by shape analysis and classification using feature descriptors like Histogram of Oriented Gradients (HOG) or Local Binary Patterns (LBP) [3]. Similarly, traditional LLR algorithms employed techniques such as Canny edge detection, color space filtering, and the Hough transform to identify and fit lines to lane markings in an image. While foundational, these methods often struggled with the immense variability of real-world driving conditions, including changing illumination, adverse weather, occlusions, and sign degradation [1].

The advent of deep learning (DL) has largely superseded these classical methods, offering end-to-end, data-driven solutions with vastly superior performance and generalization capabilities [4]. The recent timeframe that this review focuses on can be characterized not as a period of initial discovery but as one of maturation, specialization, and a critical reckoning with the persistent challenges of real-world deployment. The research community has moved beyond simply demonstrating high accuracy on clean benchmark datasets to tackling the more nuanced and difficult problems that remain [5], [6], [7].

The research goal of this review is to systematically synthesize and analyze the state-of-the-art in Traffic Sign Recognition (TSR) and Lane Line Recognition (LLR), focusing on the architectural, methodological, and conceptual shifts that have defined the field in the recent period. To structure this analysis, the paper addresses three central research questions: (1) What are the primary architectural innovations, robustness enhancement techniques, and adversarial security concerns shaping the development of modern TSR systems? (2) How is the paradigm for Lane Line Recognition shifting from 2D image-plane detection towards 3D spatial and topological understanding, and what are the key enabling technologies driving this transition? (3) What are the cross-cutting technological trends common to both TSR and LLR, and how do they influence the trajectory of the entire field of vehicle perception? The paper answers these questions through the following sections, covering advances in TSR, evolving paradigms in LLR, emerging trends, and finally synthesis and open challenges.

2. Advances in Traffic Sign Recognition (TSR)

Modern research in TSR is evolving along a clear hierarchy of abstraction. The foundational level addresses visual robustness—the challenge of 'seeing the sign clearly' by overcoming degradation from nighttime conditions, fog, or blur. The next level tackles semantic and geographical generalization, aiming to understand 'what the sign means' regardless of design variations across different countries. The highest level focuses on contextual reasoning, which requires not just recognizing the sign but determining 'how its rule applies to the road and the vehicle's actions,' a problem of multimodal integration. Crowning this hierarchy is the imperative for trust and security, which addresses adversarial robustness and explainability (XAI) to build verifiable and, ultimately, certifiable systems.

A. Architectural Innovations: The Pursuit of Speed and Accuracy

Recent research in TSR has consolidated around end-to-end models, which learn the entire pipeline from raw pixels to sign localization and classification, superseding classic multi-stage methods. The core of any modern TSR system is its neural network architecture. Recent research demonstrates a multi-pronged approach where different strategies are pursued to optimize the trade-offs between accuracy, inference speed, and model size. This development can be broadly classified into lightweight architectures for real-time deployment, high-accuracy models for maximum performance, and the integration of explainability as a core design principle.

A significant portion of research is focused on end-to-end architectures optimized for deployment on resource-constrained embedded systems found in vehicles [8]. For applications demanding real-time performance, single-pass detectors like the "You Only Look Once" (YOLO) family remain a dominant choice due to their single-pass design that prioritizes speed [9]. Recent research has focused on adapting and enhancing YOLO variants to address the specific challenges of TSR. For instance, the AF-YOLO model reconstructs the YOLOv5 backbone with a C3-SE module and an Asymptotic Feature Pyramid Network (AFPN) [10]. This redesign significantly reduces the model's parameter count by 21.79% and computational complexity by 3.73% compared to the baseline YOLOv5s, while simultaneously increasing the mean Average Precision (mAP) by 1% and boosting detection speed by 16 frames per second (FPS) [10].

Parallel to the drive for efficiency, another line of research continues to push the performance envelope. CNNs remain the bedrock of high-performance TSR systems. Studies continue to propose meticulously designed deep CNN architectures that achieve near-perfect accuracy, in some cases exceeding 99.63% on the GTSRB and 99.68% on the CTSD [2]. Models based on robust architectures like EfficientDet, which utilize advanced feature fusion such as BiFPN, have also demonstrated high performance, achieving a mAP@0.5 of 96.6% on TT100K [15]. Mirroring a broader trend in computer vision, Transformer architectures are increasingly being explored. However, their high computational cost has shifted the focus toward hybrid models that combine the strengths of CNNs and Transformers. A prime example is E-MobileViT, a lightweight model for TSR that builds upon MobileViT, an existing CNN-Vision Transformer (ViT) hybrid [16].

B. Comparative Analysis of TSR Architectural Innovations

Table I provides a comparison of recent TSR architectural innovations. Advances in traffic sign recognition showcase remarkable gains in accuracy, efficiency, and adaptability. Models like Real-time TSR CNN and E-MobileViT achieve near-perfect accuracy (ca 99.9%) on benchmarks while staying optimized for real-time or mobile deployment. AF-YOLO delivers a rare triple win—higher mAP, faster inference (345 FPS), and fewer parameters—making it ideal for embedded systems. TSD-Net sets the state-of-the-art on TT100K's small-object challenge, while YOLO-RCNN blends YOLO's speed with R-CNN's refinement for critical sign detection. Importantly, Explainable CNN adds transparency via Grad-CAM, bridging accuracy with interpretability [1], [10], [12], [16], [23].

C. Enhancing Robustness and Generalization

While achieving high accuracy on benchmark datasets is an important milestone, the true test of a TSR system is its ability to perform reliably in the unpredictable and often suboptimal conditions of real-world driving. A significant portion of recent research is therefore dedicated to enhancing model robustness against environmental adversities and challenging visual scenarios.

Environmental factors such as poor lighting, adverse weather (rain, fog), and visual obstructions are primary contributors to perception system failures. Researchers are tackling these issues from multiple angles. One of the most fundamental strategies is sophisticated data augmentation—simulating a wide range of real-world degradations such as noise, occlusions, and different weather and lighting conditions [20]. Additionally, for challenges that cannot be solved by a single sensor, multi-modal systems offer a promising solution, as demonstrated by a study using an NVIDIA Jetson Nano platform with a dual-camera setup (CSI camera for daytime and IR for nighttime) [20]. Another approach focuses on super-resolution reconstruction algorithms to enhance dim or blurred signs before recognition [21].

The detection and recognition of signs that are small (due to distance) or partially occluded remain persistent and difficult challenges. Researchers have designed specialized network architectures such as TSD-Net, which uses dynamic convolution and a refined high-resolution detection branch [23]. Improving feature fusion across different scales through detection heads like ADFF is also critical for simultaneously detecting large, nearby signs and small, distant ones [10], [23].

D. The Security Frontier: Adversarial Attacks and Defenses

As TSR systems become more capable and widely deployed, they also become more attractive targets for malicious actors. The vulnerability of deep neural networks to adversarial attacks poses a significant safety and security risk [5]. Recent research has seen an arms race between the development of more sophisticated and physically plausible attack methods and the creation of more robust defense mechanisms.

On the attack side, researchers are investigating "naturalistic" perturbations that mimic common environmental phenomena such as shadows and light patches [35]. Physical stickers present a serious threat, with research demonstrating a "universal" attack where one or two stickers can cause misclassification of any sign type [36]. The FIGhost attack takes stealth further using fluorescent ink that is invisible under normal lighting and activated only by UV light [37].

On the defense side, strategies include architectural modifications (e.g., spatial transformer networks [35]), input data transformation (e.g., multi-resolution training that low-passes adversarial noise [38]), reconstruction-based defenses (e.g., dual-mode autoencoders that detect and repair attacks [39]), and robust training methodologies (e.g., using LoRA to efficiently patch pre-trained ViTs [40]).

3. Evolving Paradigms in Lane Line Recognition (LLR)

LLR, a cornerstone of modern ADAS and autonomous driving, has undergone a period of rapid methodological evolution. The field is moving decisively beyond traditional 2D image analysis towards more sophisticated paradigms that embrace end-to-end learning, 3D spatial awareness, and multi-modal sensor fusion. These advancements are driven by the need for greater accuracy and robustness in the face of complex road geometries and challenging environmental conditions.

A. Methodological Review of 2D Lane Detection

The task of 2D lane detection involves identifying the boundaries of driving lanes within a 2D image captured by the vehicle's camera. While the goal has remained consistent, the methods for achieving it have shifted dramatically. Classical LLR pipelines were typically multi-stage, heuristic-based processes—grayscale conversion, Gaussian blur, Canny edge detection, and Hough transform. While computationally simple, their performance is brittle and highly dependent on clear lane markings and stable lighting conditions.

The modern approach, dominated by deep learning, has largely moved towards end-to-end models that learn to detect lanes directly from raw pixel data, minimizing reliance on hand-tuned parameters and fragile intermediate steps [53]. These end-to-end models can be broadly categorized into several competing paradigms: segmentation-based approaches (pixel-level classification), anchor-based and parametric approaches (pre-defined line anchors or polynomial/Bezier curve fitting), and Transformer-based approaches.

The most significant recent development in 2D LLR is the application of Transformer architectures. The self-attention mechanism at the core of Transformers allows them to capture global context and model long-range dependencies across the entire image—a crucial advantage for lane detection that enables the model to infer the path of a lane even when partially occluded [4]. Innovative Transformer-based models include MHFS-FORMER, which introduces enhanced multi-scale hybrid features and a novel multi-reference deformable attention module designed to capture the elongated structure of lanes [4], and CCHA-Net, which proposes a cross-convolutional hybrid attention mechanism achieving a very high F1 score of 80.2% on the challenging CULane dataset [56].

B. Pushing the Boundaries: 3D and Multi-Modal Lane Detection

While 2D LLR is a mature field, its fundamental limitation is that it only provides the position of lanes in the image plane. For advanced autonomous functions, the vehicle needs to know the exact 3D spatial location of lanes in the real world [6]. This has pushed 3D lane detection to the forefront of LLR research [67].

A common approach involves transforming features from the front-view (FV) image into a Bird's-Eye-View (BEV) representation. However, this FV-to-BEV transformation can be computationally expensive and prone to errors, especially on uneven road surfaces. To circumvent these issues, an emerging powerful trend is the development of BEV-free methods that predict 3D lanes directly from FV features. Anchor3DLane++ is a state-of-the-art BEV-free method that defines lane anchors directly in 3D space [67]. Its key innovation is the Prototype-based Adaptive Anchor Generation (PAAG) module, which dynamically generates a sparse set of anchors well-aligned with the lanes in the input image.

Representing a 3D lane is a critical design decision. Many methods, seeking computational efficiency, model lanes as a sparse set of 3D points. However, recent critical analysis has revealed that this sparse-point representation can be inherently flawed—the ground truth labeling process often truncates lanes at their visible endpoints, potentially leading to errors of up to 20 meters [6]. To address this, new methods propose a patching strategy and an EP-head (Endpoint head) that explicitly predicts the truncated distance to represent the complete lane structure accurately. Furthermore, the inherent challenges of camera-only systems in depth estimation have motivated research into multi-modal fusion, combining cameras with LiDAR, as demonstrated by the FusionLane model [68].

C. Beyond Markings: Virtual Lanes and Lane-Free Navigation

Research on unmarked roads leads directly to the most complex challenge: navigation in lane-free traffic environments, such as open plazas or large intersections [71]. In these scenarios, the paradigm shifts entirely from "lane detection" to dynamic generation and navigation of "virtual lanes" [72]. This requires the vehicle to not just perceive the road, but to infer a safe and efficient path based on the behavior of other agents and road context. New behavioral models such as "Lane-Free Following" (LFF) are being developed [73], and decision-making is increasingly handled by Multi-Agent Reinforcement Learning (MARL) frameworks [74], [75].

4. Emerging Trends in Perception Architectures

While TSR and LLR are distinct tasks, their progress is driven by a shared, powerful, self-reinforcing cyclical engine that shapes the entire field. This cycle begins as more powerful specialized hardware (SoCs, FPGAs) enables the creation of more complex algorithms, such as Transformers. These algorithms, in turn, 'solve' or saturate existing benchmarks, forcing the community to create new, harder datasets. These new challenges then demand even more sophisticated algorithms and more powerful hardware, completing the cycle. A critical consequence of this loop is the creation of an ever-widening 'compute chasm' between cutting-edge research and the cost- and power-constrained reality of mass-market vehicle deployment.

A. The Drive for Efficiency: Lightweight Models for Embedded Systems

A universal requirement for both TSR and LLR systems is the ability to operate in real-time on the resource-constrained computational platforms found in vehicles [16]. This has spurred a significant and shared focus on developing lightweight, efficient models that do not sacrifice accuracy. This drive for efficiency is addressed through several convergent strategies: the use of depthwise separable convolutions, the strategic use of attention mechanisms [10], [16], Binary Neural Networks (BNNs), and multi-task learning (MTL). The YOLOP model, for example, performs traffic object detection, drivable area segmentation, and lane detection within a single, unified framework [79].

B. The Data-Centric Revolution: Datasets and Benchmarks

The pace of progress in both TSR and LLR is inextricably linked to the availability of high-quality, large-scale, and diverse datasets. A clear trend in the recent literature is a data-centric approach to research, where the creation of new, more challenging benchmarks is seen as a primary driver of innovation. These new datasets focus on capturing data from adverse conditions (FoggyLane, FoggyCULane [84]), nighttime challenges (ZND Dataset [82]), 3D perception (OpenLane-V2 [7]), regional diversity (Indian Lane Dataset [85], Bangladesh Dataset [86]), and synthetic data generation [83].

C. The Rise of Integrated and Reasoning-Based Perception

The most forward-looking trend is the move towards perception systems that are more integrated and capable of reasoning, rather than simply performing isolated pattern recognition tasks. Multi-task learning (MTL) represents a first step towards this integration. A more profound development is the application of Vision-Language Models (VLMs) to enhance vehicle perception. This moves beyond simple classification and towards genuine reasoning. For instance, the LKAlert system uses a VLM to not only perform lane detection but also to predict upcoming failures of the Lane Keeping Assist system and generate natural language explanations for the driver [94]. The introduction of reasoning capabilities signals a potential paradigm shift from perception as pattern recognition to perception as reasoning.

5. Synthesis, Open Challenges, and Future Research Directions

The recent period of research in Traffic Sign and Lane Line Recognition, spanning from late 2024 into 2025, has been one of dynamic progress and increasing sophistication. The field is maturing beyond the initial deep learning revolution and is now grappling with the nuanced and difficult problems that stand in the way of widespread, safe, and reliable deployment. The literature reveals a dual pursuit: on one hand, researchers are pushing the performance ceiling with more complex and capable models; on the other, they are working to lower the floor for deployment by developing highly efficient, lightweight architectures.

Despite the remarkable advancements, several significant challenges remain. Domain Generalization is a critical issue—models trained on one dataset often experience a significant performance drop when deployed in a different region. The Long Tail problem of rare, unforeseen corner cases remains a challenge. There is a pressing need for safety-aware evaluation metrics that better reflect real-world risk. Real-world data scarcity and annotation cost also present a significant bottleneck, especially for advanced tasks like 3D lanes.

The identified challenges point towards several promising future research directions: self-supervised and lifelong learning to overcome data scarcity; tighter integration of perception and control; generative models and world models for creating realistic synthetic data; and advancing trustworthy AI through robustness, security, and interpretability [1].

6. Conclusions

This review has systematically analyzed the state of the art in Traffic Sign and Lane Line Recognition within the recent timeframe by addressing three central research questions. Regarding the evolution of TSR systems, the review found that progress is defined by three key areas: architecturally, the field shows a multi-pronged approach with continued refinement of high-accuracy CNNs, specialization of lightweight YOLO-based models, and emergence of hybrid CNN-Transformer models; in robustness, efforts focus on advanced data augmentation and multi-modal hardware solutions; and on the security front, the field is characterized by an escalating "arms race" between increasingly stealthy physical attacks and more efficient, adaptable defenses.

In response to the paradigm shift in LLR, the analysis confirms a decisive move away from 2D image-plane detection towards 3D spatial localization and topology reasoning, enabled by Transformer-based architectures, BEV-free methods, and multi-modal fusion. Finally, concerning cross-cutting trends, the review identified three themes shaping both TSR and LLR: the drive for computational efficiency, a data-centric approach, and the move towards integrated, reasoning-based perception via VLMs. While the core tasks are well-solved under standard conditions, the current research frontier is focused on achieving robustness, security, and true 3D scene understanding, laying the groundwork for the next generation of trustworthy intelligent vehicle systems.

References

Full reference list (94+ sources) available in the original PDF on IEEE Xplore. This article is published under a CC BY 4.0 license.