2d Locomotion Control of a Flywheel Based Robot Fish Using Reinforcement Learning

Alavi, Abdullah (2026) 2d Locomotion Control of a Flywheel Based Robot Fish Using Reinforcement Learning. Masters thesis, King Fahd University of Petroleum and Minerals.

[img] PDF (Master's Thesis)
Alavi_Thesis_final.pdf - Accepted Version
Restricted to Repository staff only until 10 August 2027.
Available under License Creative Commons Public Domain Dedication.

Download (32MB)

Arabic Abstract

البيئة تتسم ذلك، إلى إضافةً متعددة. أبعادًا وتتجاوز متصلة أفعاله ومساحة روبوت حالة تكون ما غالبًا يدويًا. تحديدها ويصعب بالأشياء، والإمساك لأنها للغابة معقدة الواقع في أنها إلا الحيوانات، أو للبشر بالنسبة المهمام هذه بساطة ورغم تعقيده. من يزيد مما بالتشويش، الروبوت فيها يعمل التي على الأطروحة هذه تركز بنجاح. أدائها من ليتمكن الجوهري معناها تعلم إلى الروبوت يحتاج المهام، هذه ولحل للمفاصل. منسقة حرکات تتطلب بالأسماك خاصة أنها بمعنى محددة، أنها إلا التحكم. خوارزميات من العديد ووُضعت طويلة، لفترة الأسماك سباحة دُرست المهام. هذه إحدى للاضطرابات. مقاومة تكون لا وقد التعميم، في تفشل وقد الأمثل، المستوى دون السباحة تكون قد ذلك، على علاوة بالبيئة. ومرتبطة المدروسة الخوارزمية تكون أن يجب المثلى. السباحة استراتيجية وإيجاد الأسماك سباحة مشكلة لمعالجة المعزز التعلم استخدام إلى نهدف العمل، هذا في خوارزمية هناك تكون أن يجب أخرى، بعبارة المستوى. في مكان أي إلى والوصول السباحة من السمكة تتمكن بحيث الحركة نطاق تعميم على قادرة السمكة. من الاتجاه أو المسافة عن النظر بغض نقطة أي إلى والوصول اتجاه أي في والانعطاف للأسام السباحة مهمة إنجاز على قادرة فقط واحدة أخرى خوارزميات استخدام أيضًا سيتم .(ARS) المعزز العشوائي البحث باسم المهمة هذه لتعلم المستخدمة الأساسية المعزز التعلم خوارزمية تُعرف حركة تحقيق الممكن من أنه ARS خوارزمية باستخدام الأولية المحاكاة نتائج تُظهر أداءً. النماذج أفضل وإيجاد للمقارنة DQN و A2C و PPO مثل السمكة توجيه تم التقييم، مرحلة في فقط. واحد اتجاه في مستقيم بشكل السباحة على السمكة تدريب تم للتوضيح، بعدين. في معممة سباحة زاوية. سرعة بأي والانعطاف اتجاه أي في السباحة تستطيع أنها النتائج تُظهر عليها. تدريبها يتم لم أخرى زاوية وبسرعات أخرى اتجاهات في للسباحة متعددة خارجية استشعار أجهزة أو متعددة، خوارزميات تستخدم ما غالبًا أنها إلا مماثلة، أهدافًا حققت قد المجال هذا في السابقة الدراسات أن رغم الرئيسية المساهمة تكمن تنفيذه. وتكلفة النظام تعقيد من يزيد هذا كل طويلة. تدريب فترات أو ،(GPS) العالمي المواقع تحديد نظام أو كالكاميرات طويلة. تدريب فترات أو متعددة استشعار أجهزة أو خوارزميات استخدام دون فعالة عامة سباحة تحقيق إمكانية وإثبات تبسيط في الدراسة لهذه

English Abstract

Reinforcement learning has emerged as an essential tool used to teach a robot complex skills or tasks. The tasks may range from walking or running to grasping or holding objects and are extremely hard to define manually. Often the state and action space of the robot are continuous and exceeds many dimensions. In addition, the environment they operate in are quite noisy. This adds additional layer of complexity. Though such tasks are very simple for human beings or animals, they are in fact quite complex as they require coordinated movements by the joints. To solve such tasks, the robot needs to learn the intrinsic meaning of such tasks in order to perform such tasks without failing. The focus of this thesis is among such tasks. Fish swimming has been studied for a long time. A lot of control algorithms have been defined. However, they are specific in the sense that they are specific for the fish in study and deterministic to the environment. In addition, the swimming gait may be suboptimal, fail to attain generalization, and may not be robust to disturbances. In this work, we aim to use reinforcement learning to tackle the problem of fish swimming and find the optimal swimming strategy. The algorithm should also be able to generalize a range of motion so that the fish can swim and reach anywhere on the plane. In other words, there should be only one algorithm that can fulfill the task of swimming forward and turning to any direction and reach any arbitrary point no matter the distance or direction from the fish. The primary reinforcement learning algorithm used to learn this task is known as Augmented Random Search (ARS). Other algorithms like PPO, A2C, DQN, and will also be used to compare and find the best performing models. The preliminary simulation results using the ARS algorithm shows that it is possible to achieve generalized swimming motion in two dimensions. To elaborate, the fish was trained to swim straight only in a single direction. In the evaluation phase, the fish was instructed to swim in other directions and other angular velocities that it was not trained for. The results show that it can swim in any direction and turn with any angular velocity. Although existing works in the literature have achieved similar goals, they either often use multiple algorithms, or multiple external sensors like cameras or GPS, or long training times. All of this increases the complexity of the system and the cost of implementation. The key contribution of this work is to to simplify and show that it is possible to achieve efficient generalized swimming without the use of multiple algorithms, sensors or long training times.

Item Type: Thesis (Masters)
Subjects: Engineering
Electrical
Department: College of Engineering and Physics > Electrical Engineering
Thesis Advisor:
Ali Al-beladi,
Thesis Committee Members:
Mudassir Masood Ali, Brahim Brahmi,
Depositing User: ABDULLAH ALAVI
Date Deposited: 11 Aug 2026 05:26
Last Modified: 11 Aug 2026 05:26
URI: https://eprints.kfupm.edu.sa/id/eprint/144690