Skip to main content

Data Mining

 Este estudio sintetizó cinco años de datos del comportamiento de estudiantes en 27 variables relacionadas con las dimensiones socioeconómicos, académicos y familiares, hasta generar un dataset genérico replicable para universidades públicas y privadas en relación al abandono. Para este estudio fue posible procesar 3.006.531 registros hasta obtener información coherente de 13.715 estudiantes 

El dataset desarrollado integra múltiples dimensiones relevantes para el análisis académico y profesional de los estudiantes.

Categories:

Please cite the following paper when using this dataset:

Vanessa Su and Nirmalya Thakur, “COVID-19 on YouTube: A Data-Driven Analysis of Sentiment, Toxicity, and Content Recommendations”, Proceedings of the IEEE 15th Annual Computing and Communication Workshop and Conference 2025, Las Vegas, USA, Jan 06-08, 2025 (Paper accepted for publication, Preprint: https://arxiv.org/abs/2412.17180).

Abstract:

Categories:

To download the dataset without purchasing an IEEE Dataport subscription, please visit: https://zenodo.org/records/13738598

Please cite the following paper when using this dataset:

N. Thakur, “Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,” arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292

Abstract

Categories:

Recently, combinatorial interaction strategies have a large spectrum as black box strategies for testing software and hardware. This paper discusses a novel adoption of a combinatorial interaction strategy to generate a sparse combinatorial data table (SCDT) for machine learning. Unlike test data generation strategies, in which the t-way tuples synthesize into a test case, the proposed SCDT requires analyzing instances against their corresponding tuples to generate a systematic learning dataset.

Categories:

Please cite the following paper when using this dataset:

N. Thakur, V. Su, M. Shao, K. Patel, H. Jeong, V. Knieling, and A.Bian “A labelled dataset for sentiment analysis of videos on YouTube, TikTok, and other sources about the 2024 outbreak of measles,” Proceedings of the 26th International Conference on Human-Computer Interaction (HCII 2024), Washington, USA, 29 June - 4 July 2024. (URL: https://dl.acm.org/doi/10.1007/978-3-031-76806-4_17)

Abstract

Categories:

The alignment between the implemented database application systems and their data specification standard description files significantly affects the accuracy of enterprises' estimation of data assets based on data standard files. In this study, we proposed an automated approach for discovering and aligning these consistent fields, greatly reducing the cost of manual evaluation. We frame the field's alignment problem as an entity matching computation on two distinct graphs, respectively constructed from the database of application systems and its data specification standard.

Categories:

The fast development of urban advancement in the past decade requires reasonable and realistic solutions for transport, building infrastructure, natural conditions, and personal satisfaction in smart cities. This paper presents and explores predictive energy consumption models based on data-mining techniques for a smart small-scale steel industry in South Korea. Energy consumption data is collected using IoT based systems and used for prediction.

Categories:

India is known for its highly disciplined foreign policies, strategic location, vibrant and massive Diaspora. India envisages enhancing its scope of cooperation, trade and widens its sphere of relations with the Pacific. As a result, the world is witnessing the rise of Indo-Pacific ties. Before the 1980’s the keystone of the universe was called the Atlantic, but now a radical shift to the east is noticed by the term “Indo-Pacific‟.

Categories: