Research Memo

메모: Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring

Fig.1 Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring

Published in: IEEE Transactions on Artificial Intelligence ( Volume: 5, Issue: 12, December 2024)

Date of Publication: 05 September 2024

Memo

앞서 확인한 다른 논문과는 달리 컬러 이미지를 사용했고, frame difference에 중점을 두기보다는 attention mechanism을 추가해 주요 기여는 이쪽이 아닐까 싶다.

후자는 내가 사용하기에는 NN에 대한 이해가 미천하다.

Introduction

뭔가 abstract의 convention은 안 따르긴 하는데, 무튼 UAV로 찍은 6600장의 데이터로, YOLO 기반 obj. detection을 개선해 16.22%의 성능 개선을 보였다고 함.

주 기여 3가지:

  1. 데이터셋 만듦
  2. YOLOv5에 spatiotemporal data를 도입하는 방안 연구
  3. Attention mechanism에 spatiotemporal data 도입이 성능 향상을 보이는 것을 확인

생략

Body

Method

4가지 기법을 제시하고 비교함.

  1. 기본 YOLOv5에 프레임 하나만 넣음
  2. 기본 YOLOv5에 프레임 둘을(t, t+1) concat하여 6-channel image로 넣음
  3. 기본 YOLOv5에 프레임 둘(t, t+1), greyscale 후 frame difference를 수행한 데이터까지 concat하여 7-channel image 넣음
  4. YOLOv5의 backbone을 두 개로 만들어, 각각 2번의 6-channel data와, 3의 1-channel frame difference data를 넣고, 이 each backbone의 결과를 concat하여 head의 input으로 넣는 형태(정확히는 각 backbone의 layer에서 나오는 것도 concat하여 넣는다. 원문 fig. 7 확인)

여기에 3,4번 방법에는 attention mechanism을 추가하는데, 이것도 두 가지 방법으로 나뉜다. 하나는 backbone의 마지막 부분, 즉 최종 spatial pyramid pooling layer 직전에 하나의 layer로 추가하는 것이고, 다른 하나는 head architecture의 서로 다른 네 개 level에 위치한 C3 module 내부에 attention mechanism을 embedding하는 것.

Exp.

결과는 앞서 제시한 1-4 방법과 두 가지 attention mechanism을 교차해 P, R, mAP50을 확인하였다.

Conclusion


참고문헌

K. Telegraph and C. Kyrkou, “Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring,” in IEEE Transactions on Artificial Intelligence, vol. 5, no. 12, pp. 6159-6171, Dec. 2024, doi: 10.1109/TAI.2024.3454566.

Leave a comment