Publications | https://photogrammetry.survey.ntua.gr

Boutsi, A-M; Bakalos, N; Ioannidis, C

POSE ESTIMATION THROUGH MASK-R CNN AND VSLAM IN LARGE-SCALE OUTDOORS AUGMENTED REALITY Journal Article

In: ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci., vol. V-4-2022, no. 4, pp. 197–204, 2022, ISSN: 2194-9050.

Abstract | Links | BibTeX | Tags: 3D rendering, Augmented Reality, CNN, deep learning, image recognition, pose estimation

@article{boutsi2022pose,

title = {POSE ESTIMATION THROUGH MASK-R CNN AND VSLAM IN LARGE-SCALE OUTDOORS AUGMENTED REALITY},

author = {A-M Boutsi and N Bakalos and C Ioannidis},

url = {https://www.isprs-ann-photogramm-remote-sens-spatial-inf-sci.net/V-4-2022/197/2022/},

doi = {10.5194/isprs-annals-V-4-2022-197-2022},

issn = {2194-9050},

year  = {2022},

date = {2022-05-18},

journal = {ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci.},

volume = {V-4-2022},

number = {4},

pages = {197--204},

abstract = {Abstract. Deep Learning (DL) ingrained into Mobile Augmented Reality (MAR) enables a new information-delivery paradigm. In the context of 6 DoF pose estimation, powerful DL networks could provide a direct solution for AR systems. However, their concurrent operation requires a significant number of computations per frame and yields to both misclassifications and localization errors. In this paper, a hybrid and lightweight solution on 3D tracking of arbitrary geometry for outdoor MAR scenarios is presented. The camera pose information obtained by ARCore SDK and vSLAM algorithm is combined with the semantic and geometric output of a CNN-object detector to validate and improve tracking performance in large-scale and uncontrolled outdoor environments. The methodology involves three main steps: i) training of the Mask-R CNN model to extract the class, bounding box and mask predictions, ii) real-time detection, segmentation and localization of the region of interest (ROI) in camera frames, and iii) computation of 2D-3D correspondences to enhance pose estimation of a 3D overlay. The dataset holds 30 images of the rock of St. Modestos \textendash Modi in Meteora, Greece in which the ROI is an area with characteristic geological features. The comparative evaluation between the prototype system and the original one, as well as with R-CNN and FAST-R CNN detectors demonstrates higher precision accuracy and stable visualization at half a kilometre distance, while tracking time has decreased at 42% during far-field AR session.},

keywords = {3D rendering, Augmented Reality, CNN, deep learning, image recognition, pose estimation},

pubstate = {published},

tppubtype = {article}

}

Close

Bakalos, Nikolaos; Rallis, Ioannis; Doulamis, Nikolaos; Doulamis, Anastasios; Voulodimos, Athanasios; Protopapadakis, Eftychios

Adaptive Convolutionally Enchanced Bi-Directional Lstm Networks For Choreographic Modeling Proceedings Article

In: 2020 IEEE Int. Conf. Image Process., pp. 1826–1830, IEEE, 2020, ISBN: 978-1-7281-6395-6.

Abstract | Links | BibTeX | Tags: CNN, Convolutional LSTM, Folkloric dances, Intangible Cultural Heritage, LSTM, Posture identification

Agrafiotis, P; Skarlatos, D; Forbes, T; Poullis, C; Skamantzari, M; Georgopoulos, A

Underwater photogrammetry in very shallow waters: Main challenges and caustics effect removal Journal Article

In: Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. - ISPRS Arch., vol. 42, no. 2, pp. 15–22, 2018, ISSN: 16821750.

Abstract | Links | BibTeX | Tags: Caustics, CNN, SfM MVS, Underwater 3D reconstruction