
Would you believe it if I tell you that you are a part of one of the sources generating big data? A layperson may instantly disagree, but it is, in fact, valid for most people using smartphones. Whether you are using social media, sending money to someone online, or using a variety of mobile applications, all of your web activity leads to data generation. This data is precious to companies who want to understand their customers better and provide personalized services. For example, eCommerce giant Amazon uses your shopping history and past product search data to recommend you related products you are most likely to buy. Research shows that we generate around 2.5 quintillion bytes of data each day, which is quite huge! Well, you may wonder how can such a massive amount of data be handled. This is where Hadoop comes into the picture. Apache Hadoop is one of the earliest open-source tools to offer storage and large-scale processing of big data. As mentioned in Apache.org, Hadoop is a framework that allows for distributed processing of large data sets across clusters of computers using simple programming models. Various companies and organizations use Hadoop for research as well as production.
Continue reading