1,863 results for Recognition · 0.119s

Sponsored Partners
arxiv.org/abs/2406.12931v1

Automatic Speech Recognition for Biomedical Data in Bengali Language

This paper presents the development of a prototype Automatic Speech Recognition (ASR) system specifically designed for Bengali biomedical data. Recent advancements in Bengali ASR are encouraging, but a lack of domain-specific data limits the creation...

arxiv.org/abs/2207.10024v1

Difficulty-Aware Simulator for Open Set Recognition

Open set recognition (OSR) assumes unknown instances appear out of the blue at the inference time. The main challenge of OSR is that the response of models for unknowns is totally unpredictable. Furthermore, the diversity of open set makes it harder...

arxiv.org/abs/2204.06328v2

HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition

Pre-training with self-supervised models, such as Hidden-unit BERT (HuBERT) and wav2vec 2.0, has brought significant improvements in automatic speech recognition (ASR). However, these models usually require an expensive computational cost to achieve...

arxiv.org/abs/2204.04564v1

Multimodal Transformer for Nursing Activity Recognition

In an aging population, elderly patient safety is a primary concern at hospitals and nursing homes, which demands for increased nurse care. By performing nurse activity recognition, we can not only make sure that all patients get an equal desired car...

arxiv.org/abs/2102.04652v1

Large Scale Long-tailed Product Recognition System at Alibaba

A practical large scale product recognition system suffers from the phenomenon of long-tailed imbalanced training data under the E-commercial circumstance at Alibaba. Besides product images at Alibaba, plenty of image related side information (e.g. t...

arxiv.org/abs/2508.05055v1

MOVER: Combining Multiple Meeting Recognition Systems

In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to combine the output of diarization (e.g., DOVER) or automatic speech rec...

arxiv.org/abs/1801.06867v1

Scene recognition with CNNs: objects, scales and dataset bias

Since scenes are composed in part of objects, accurate recognition of scenes requires knowledge about both scenes and objects. In this paper we address two related problems: 1) scale induced dataset bias in multi-scale convolutional neural network (C...

arxiv.org/abs/2104.03419v1

Towards On-Device Face Recognition in Body-worn Cameras

Face recognition technology related to recognizing identities is widely adopted in intelligence gathering, law enforcement, surveillance, and consumer applications. Recently, this technology has been ported to smartphones and body-worn cameras (BWC)....