Machine Learning

Message Queuing Telemetry Transport (MQTT) protocol is one of the most used standards used in Internet of Things (IoT) machine to machine communication. The increase in the number of available IoT devices and used protocols reinforce the need for new and robust Intrusion Detection Systems (IDS). However, building IoT IDS requires the availability of datasets to process, train and evaluate these models. The dataset presented in this paper is the first to simulate an MQTT-based network. The dataset is generated using a simulated MQTT network architecture.

Categories:
21958 Views

Invasive lobular carcinoma (ILC) is the second most prevalent histologic subtype of invasive breast cancer. Here, we comprehensively profiled 817 breast tumors, including 127 ILC, 490 ductal (IDC), and 88 mixed IDC/ILC. Besides E-cadherin loss, the best known ILC genetic hallmark, we identified mutations targeting PTEN, TBX3 and FOXA1 as ILC enriched features. PTEN loss associated with increased AKT phosphorylation, which was highest in ILC among all breast cancer subtypes. Spatially clustered FOXA1 mutations correlated with increased FOXA1 expression and activity.

Categories:
616 Views

This dataset is a large-scale Chinese hotel review data set collected by Tan Songbo.  The corpus size is 10,000 reviews. The corpus is automatically collected and organized from Trip.com.

Categories:
1155 Views

This dataset was created from all Landsat-8 images from South America in the year 2018. More than 31 thousand images were processed (15 TB of data), and approximately on half of them active fire pixels were found. The Landsat-8 sensor has 30 meters of spatial resolution (1 panchromatic band of 15m), 16 bits of radiometric resolution and 16 days of temporal resolution (revisit). The images in our dataset are in TIFF (geotiff) format with 10 bands (excluding the 15m panchromatic band).

Categories:
6217 Views

The dataset includes 2 parts: private and public traffic.

The private traffic is self-captured network traffic of serveral softwares, such as YouTube, Skype, streaming video, totally 16 categories.

The public traffic is an open VPN dataset, including numorous VPN or nonVPN network services, totally 24 categories.

 

Categories:
1130 Views

The dataset contains rash images of 11 different disease states. Images of normal skin are also included in the dataset.

Categories:
17692 Views

Spoken Indian Language Identification Database

(9 languages, 8 different utterance lengths)

Languages

  1. Assamese 
  2. Bengali 
  3. Gujarati 
  4. Hindi 
  5. Kannada 
  6. Malayalam 
  7. Marathi 
  8. Tamil 
  9. Telugu

Durations

  1. 30 sec
  2. 10 sec
  3. 5 sec
  4. 3 sec
  5. 1 sec
  6. 0.5 sec
  7. 0.2 sec
  8. 0.1 sec

 

 

Categories:
1108 Views

This dataset extends the Urban Semantic 3D (US3D) dataset developed and first released for the 2019 IEEE GRSS Data Fusion Contest (DFC19). We provide additional geographic tiles to supplement the DFC19 training data and also new data for each tile to enable training and validation of models to predict geocentric pose, defined as an object's height above ground and orientation with respect to gravity. We also add to the DFC19 data from Jacksonville, Florida and Omaha, Nebraska with new geographic tiles from Atlanta, Georgia.

Categories:
10066 Views

The raw data are collected from the websites of EPD (Environmental Protection Department, Hong Kong) and HKO (Hong Kong Observatory). Marine water quality data is provided by EPD and climatological data is provided by HKO. The data is interpolated by SAS “proc expand” and aligned to the beginning of each month.

 

The raw data used to produce this dataset are extracted from the following URL.

Categories:
1322 Views

Cautionary traffic signs are of immense significance to traffic safety. In this study,  a robust and optimal real-time approach to recognize the Indian Cautionary Traffic Signs(ICTS) is proposed. ICTS are all triangles with a white backdrop, a red border, and a black pattern. A dataset of 34,000 real-time images has been acquired under various environmental conditions and categorized into 40 distinct classes. Pre-processing techniques are used to transform RGB images to Gray-scale images and enhance contrast in images for superior performance.

Categories:
9832 Views

Pages