The exponential growth of data necessitates efficient solutions for information processing, especially for iterative algorithms used in Machine Learning. While platforms like Hadoop with MapReduce are effective for massive processing, they present challenges with iterative workloads. This project addresses this limitation by proposing the design of a cluster computing architecture based on Apache Spark. Spark, an open-source platform, is ideal due to its fault tolerance and its ability to accelerate iterative algorithms through in-memory processing across the cluster. The main objective is to design a robust architecture, deployable locally or in the cloud, that maximizes the performance of advanced Machine Learning techniques (such as Neural Networks or Bayesian Networks) which demand high processing capacity and low latency. The methodology includes a descriptive phase for the state-of-the-art review, followed by an architectural design phase, an experimental phase for objective validation, and will culminate in the publication of the results in a JCR scientific journal.<br/><br/><b>Goal</b>: <br/>Design an optimized cluster computing architecture for efficiently executing Machine Learning techniques, leveraging the capabilities of the Apache Spark distributed processing engine.<br/><br/><b>Research lines</b>: <br/>Artificial intelligence and data mining
| Status | Finished |
|---|
| Effective start/end date | 2/04/18 → 2/04/19 |
|---|
In 2015, UN member states agreed to 17 global Sustainable Development Goals (SDGs) to end poverty, protect the planet and ensure prosperity for all. This project contributes towards the following SDG(s):
-
SDG 7
Affordable and Clean Energy