Skip to main navigation Skip to search Skip to main content

Machine Learning Techniques Using Apache Spark

Project Details

Description

The exponential growth of data necessitates efficient solutions for information processing, especially for iterative algorithms used in Machine Learning. While platforms like Hadoop with MapReduce are effective for massive processing, they present challenges with iterative workloads. This project addresses this limitation by proposing the design of a cluster computing architecture based on Apache Spark. Spark, an open-source platform, is ideal due to its fault tolerance and its ability to accelerate iterative algorithms through in-memory processing across the cluster. The main objective is to design a robust architecture, deployable locally or in the cloud, that maximizes the performance of advanced Machine Learning techniques (such as Neural Networks or Bayesian Networks) which demand high processing capacity and low latency. The methodology includes a descriptive phase for the state-of-the-art review, followed by an architectural design phase, an experimental phase for objective validation, and will culminate in the publication of the results in a JCR scientific journal.<br/><br/><b>Goal</b>: <br/>Design an optimized cluster computing architecture for efficiently executing Machine Learning techniques, leveraging the capabilities of the Apache Spark distributed processing engine.<br/><br/><b>Research lines</b>: <br/>Artificial intelligence and data mining
StatusFinished
Effective start/end date2/04/182/04/19

UN Sustainable Development Goals

In 2015, UN member states agreed to 17 global Sustainable Development Goals (SDGs) to end poverty, protect the planet and ensure prosperity for all. This project contributes towards the following SDG(s):

  1. SDG 7 - Affordable and Clean Energy
    SDG 7 Affordable and Clean Energy

Keywords

  • Cluster Computing
  • Machine Learning
  • Distributed Processing
  • Apache Spark
  • Iterative Algorithms
  • System Architecture
  • Big Data
  • Hadoop
  • MapReduce
  • Neural Networks

CACES Knowledge Areas

  • 116A Computer Science

Categorías UNESCO

  • Software and application development and analysis

Fingerprint

Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint.