Identifikasi Objek Motor dengan Vision Transformer pada Dataset Caltech 101

Authors

  • Refianto Damai Darmawan Universitas Muhammadiyah Bogor Raya
  • Ruhiat Susanto Universitas Muhammadiyah Bogor Raya

Keywords:

deteksi objek, bounding box, identifikasi, pengolahan citra digital, deep learning

Abstract

Penelitian ini bertujuan untuk mengidentifikasi objek motor menggunakan arsitektur Vision Transformer (ViT) pada dataset Caltech 101, dengan fokus pada kategori "motorcycles" untuk meningkatkan akurasi deteksi bounding box. Di era digital, visi komputer memainkan peran penting dalam keselamatan lalu lintas, terutama di Indonesia di mana sepeda motor mendominasi dengan lebih dari 139 juta unit menurut Badan Pusat Statistik (2024). Dataset Caltech 101, yang terdiri dari 9.146 gambar dalam 101 kategori, dipilih karena keberagamannya meskipun ukurannya relatif kecil, sehingga ideal untuk pengujian model pada data terbatas. Metodologi mencakup akuisisi data (798 gambar motorcycles dengan anotasi bounding box), praproses (resizing ke 224x224 piksel dan normalisasi), serta pelatihan ViT dengan 4 lapisan encoder, optimizer AdamW (learning rate 1e-3), dan early stopping. Dataset dibagi 80% pelatihan dan 20% pengujian. Evaluasi menggunakan Intersection over Union (IoU) sebagai metrik utama. Hasil menunjukkan rata-rata IoU 83,33%, dengan variasi dari 51% hingga 95% pada sampel, menandakan performa baik meskipun tanpa bobot pre-trained. Pembahasan menyoroti keunggulan ViT dalam menangkap relasi global, tetapi juga keterbatasan seperti kesalahan anotasi dataset, bias objek negara Barat, dan kebutuhan komputasi tinggi. Aplikasi praktis meliputi integrasi dengan CCTV untuk pemantauan lalu lintas cerdas di kota pintar. Kesimpulan: ViT potensial untuk deteksi objek motor, meskipun memerlukan perbaikan dataset dan adaptasi untuk konteks nyata seperti lalu lintas Indonesia. Penelitian ini memperkaya literatur visi komputer dan membuka peluang inovasi transportasi berkelanjutan.

References

Badan Pusat Statistik (2024). Perkembangan Jumlah Kendaraan Bermotor Menurut Jenis (Unit), 2024 [Laporan Statistik]. Diakses dari http://bps.go.id (Tanggal akses 8 Desember 2025).

Li, F.-F., Andreeto, M., Ranzato, M., & Perona, P. (2022). Caltech 101 (1.0) [Data set]. CaltechDATA. doi:10.22002/D1.20086.

Permanasari, Y., Ruchjana, B. N., Hadi, S., & Rejito, J. (2022). Innovative region convolutional neural network algorithm for object identification. Journal of Open Innovation: Technology, Market, and Complexity, 8(4), 182. doi:10.3390/joitmc8040182

Montserrat, D. M., Lin, Q., Allebach, J., & Delp, E. J. (2017). Training object detection and recognition CNN models using data augmentation. Electronic Imaging, 29, 27-36. doi: 10.2352/ISSN.2470-1173.2017.10.IMAWM-163

Verma, N. K., Sharma, T., Rajurkar, S. D., & Salour, A. (2016, October). Object identification for inventory management using convolutional neural network. In 2016 IEEE Applied Imagery Pattern Recognition Workshop (AIPR) (pp. 1-6). IEEE. doi: 10.1109/AIPR.2016.8010578

Gillioz, A., Casas, J., Mugellini, E., & Abou Khaled, O. (2020, September). Overview of the Transformer-based Models for NLP Tasks. In 2020 15th Conference on computer science and information systems (FedCSIS) (pp. 179-183). IEEE. doi: 10.15439/2020F20

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.

Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., & Beyer, L. (2021). How to train your vit? data, augmentation, and regularization in vision transformers. arXiv preprint arXiv:2106.10270.

Zhu, S., Yang, T., & Chen, C. (2021). Visual explanation for deep metric learning. IEEE Transactions on Image Processing, 30, 7593-7607. doi: 10.1109/TIP.2021.3107214

Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., ... & Zitnick, C. L. (2014, September). Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Cham: Springer International Publishing. doi:10.1007/978-3-319-10602-1_48

Xiong, Y., Lan, L. C., Chen, X., Wang, R., & Hsieh, C. J. (2022, January). Learning to schedule learning rate with graph neural networks. In International conference on learning representation (ICLR).

Chu, C., Zhmoginov, A., & Sandler, M. (2017). Cyclegan, a master of steganography. arXiv preprint arXiv:1712.02950.

Nguyen, D., Hoang, V. D., & Le, V. T. L. (2024, April). V-DETR: Pure Transformer for End-to-End Object Detection. In Asian Conference on Intelligent Information and Database Systems (pp. 120-131). Singapore: Springer Nature Singapore. doi:10.1007/978-981-97-4985-0_10

Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009, June). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255). Ieee. doi:10.1109/CVPR.2009.5206848

Hendrycks, D., Lee, K., & Mazeika, M. (2019, May). Using pre-training can improve model robustness and uncertainty. In International conference on machine learning (pp. 2712-2721). PMLR.

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25. doi:

Published

2025-10-30

How to Cite

Darmawan, R. D., & Susanto, R. (2025). Identifikasi Objek Motor dengan Vision Transformer pada Dataset Caltech 101. JINTIKOM : Jurnal Informasi Teknologi Dan Komputer , 1(2), 55–67. Retrieved from https://journal.umbogorraya.ac.id/index.php/jintikom/article/view/387