publications
research topics I have worked upon..
2026
- EAAIAcceptedMultimodal Exponential Moving Average-guided Representation for Integrated Tumor-segmentation and Survival PredictionAshis Datta, Shashwat Sarkar, Aaditya Lochan Sharma, Palash GhosalIn Engineering Applications of Artificial Intelligence, 2026
Glioblastoma survival prediction is a vital challenge in brain tumor research, yet most existing approaches still use isolated or staged pipelines that separate tumor segmentation, feature design, and survival prediction. These staged setups often carry errors from one step to the next resulting in error accumulation, it depends on time-consuming handcrafted features, and fails to link early outputs directly to patient survival. To overcome these issues, we introduce a student–teacher framework that combines volumetric tumor segmentation and survival prediction through knowledge distillation. We introduce a segmentation framework which builds on Shifted Windows UNet Transformer (Swin-UNETR) backbone with attention modules and multi-scale pooling to segment detailed tumor structure from Magnetic Resonance Imaging (MRI) scans. We then use a multimodal Transformer to merge these imaging features with segmentation maps and patient data such as age and surgical resection. By employing an Exponential Moving Average (EMA) as the teacher of the Multimodal Transformer student, our approach transfers a stable risk signal while learning image patterns that connect with clinical outcomes. This student–teacher training eliminates the need for handcrafted radiomics and avoids the issues observed in staged pipelines. Overall, our student–teacher design frames survival prediction as a knowledge distillation problem, offering a more direct and meaningful path from raw scans to patient survival prediction.
2025
- ICDAR WMLAcceptedA New Multimodal Cross-Domain Network for Classification of Challenging Scene ImagesShashwat Sarkar, Kunal Purkayastha, Shivakumara Palaiahnakote, Muhammad Hammad Saleem, Umapada Pal, Palash GhosalIn ICDAR Workshop on Machine Learning, 2025
Classification of natural-scene images affected by adverse effects into their respective challenge type is a challenging problem. The state-of-the-art usually focuses on developing a model that extracts invariant features to address the above challenges. However, this direction is ineffective and inefficient for solving industry-oriented problems. This is because industry-related problems are specific and have a limited dataset and complexity. Therefore, this work introduces a new multi-modal approach called data-centric, which focuses on the classification of images of challenges (complexities), namely, quality degradation, illumination effects, and contrast issues. Once the complexity is known, the existing method can be used to achieve the expected results without proposing a new model. Our work proposes a dual ResNet50 for extracting spatial-domain and discrete wavelet-domain based features, which outputs image features. To extract the features that assess the quality of the images, the proposed work introduces a dual-CLIP-based text encoder by considering class labels as input. Therefore, the proposed model leverages image and textual features for the classification. The approach integrates image features in both spatial and wavelet domains, along with the textual context of the image quality, such as defect class labels. To evaluate the proposed method, the classification rate is estimated on our dataset and compared with the state-of-the-art methods.
- SNCSAcceptedA Template Matching Based Approach for Geolocating Cadastral Aerial ImagesPraveen Kumar Pradhan, Aaditya Lochan Sharma, Shashwat Sarkar, Udayan Baruah, Biswaraj Sen, Palash GhosalIn SN Computer Science, 2025
The rapid increase in aerial image analysis has created a compelling need to accurately determine the geolocation of regions of small, cropped images from a large context aerial image. The challenge lies in matching any land area captured from a drone or satellite to their corresponding locations within larger, full-scale cadastral maps. This is a task complicated by variations in scale, resolution, and land cover characteristics. To address this problem, an aerial image segmentation dataset and a cadastral map dataset were introduced by first segmenting cropped aerial images into distinct land cover classes such as Buildings, Roads, Vegetation, Trees, and Barren areas. These segmentations are then transformed into simplified cadastral maps highlighting critical boundaries using a deep learning-based multi-class segmentation method. Consecutively, we process the full-scale cadastral map along with randomly selected patches through a novel Deep Feature Template Matching Network (DFTMN). This network is designed to learn the correspondence between local image patches and their global context within the larger map, effectively aligning fine-grained features with their broader spatial surroundings. The proposed framework demonstrates strong potential for applications in urban planning, environmental monitoring, and disaster response, offering an effective solution for fine-to-coarse scale geolocation in aerial imagery.
2024
- ICPRAcceptedDATR: Domain Agnostic Text RecognizerKunal Purkayastha, Shashwat Sarkar, Umapada Pal, Shivakumara Palaiahnakote, Palash GhosalIn International Conference on Pattern Recognition, 2024
Recognizing text extracted from multiple domains is complex and challenging because complexities vary from one domain to another. Most existing methods focus either on natural scene text or specific text type but not text of multiple domains, namely, scene, underwater, and drone texts. In addition, the state-of-the-art models ignore the vital cues that exist in multiple instances of the text. This paper presents a new method called the Student-Teacher-Assistant (STA) network, which involves dual CLIP models to exploit cues in multiple text instances. The model that uses ResNet50 in its image encoder is called helper CLIP, while the model that uses ViT in its image encoder is called primary CLIP. The proposed work processes both models simultaneously to extract visual and textual features through image and text encoders. Our work uses cosine similarity for the randomly chosen input image to detect instances similar to the input image. The input and similar instances are supplied to primary and helper CLIPs for visual and textual feature extraction. The outputs of dual CLIPs are fused in a different way through the alignment step for recognizing text accurately, irrespective of domains. To demonstrate the proposed model’s significance, experiments are conducted on a set of standard natural scene text datasets (regular and irregular), underwater images, and drone images. The results on three different domains show that the proposed model outperforms the state-of-the-art recognition models.
- ICPRAcceptedDITS: Domain-Independent Text SpotterKunal Purkayastha, Shashwat Sarkar, Umapada Pal, Palaiahnakote Shivakumara, Palash GhosalIn International Conference on Pattern Recognition, 2024
Text spotting in diverse domains, such as drone-captured images, underwater scenes, and natural scene images, presents unique challenges due to variations in image quality, contrast, text appearance, background complexity, and external factors like water surface reflections and weather conditions. While most existing approaches focus on text spotting in natural scene images, we propose a Domain-Independent Text Spotter (DITS) that effectively handles multiple domains. We innovatively combine the Real-ESRGAN, developed for regular image enhancement, with the DeepSolo, developed for scene text spotting, in an end-to-end fashion for text detection and spotting on images of different domains. The key idea behind our approach is that improving image quality and text-spotting accuracy are complementary goals. Real-ESRGAN enhances image quality, making the text more discernible, while DeepSolo, a state-of-the-art text spotting model, accurately localizes and recognizes text in the enhanced images. We validate the superiority of our proposed model by evaluating it on datasets from drone, underwater, and scene domains (ICDAR 2015, CTW1500, and Total-Text).