Abstract
Log data refers to a detailed list of application information, system performance, or user activities in a system with their respective timestamp. Logs are useful for keeping track of customer activities and understanding about the usage of the system. Logs can also be used for monitoring and ensuring the proper working of the system which can be reviewed chronologically. Systems usually store logs in MySQL databases and perform queries over them. They hugely rely on batch processing which does not support real-time search and analytics. Operations over the logs should be performed on real-time so that decisions and other improvements can be carried out as early as possible. Our proposed system ensures proper handling of logs, make them available for real-time search and store them in distributed environment. Proposed method uses Apache Kafka as message broker, Apache Hadoop HDFS for archival of log data and Elastic search to power near real-time search. Proposed system allows multiple clients to query log data, perform metrics and grouping analytics on them. We also make sure that logs are passed to the correct consumer through our message broker. Analytics over the log data helps us to visualize the pattern of usage of the system and its features with which we decide the performance improvements and business decisions. Further these can be represented as charts, graphs or tables for better visualization.
Keywords
Cloud Computing
Data Indexing
Scalable Architecture
Information Retrieval
Big Data Analytics
Authors
How to Cite this Article
Maria Michael Visuwasam, Abijith Krishna.U, Balakumaran.M (2016).
"CLOUD BASED LOG STORAGE AND ARCHIVAL WITH REAL TIME SEARCH CAPABILITY".
International Journal of Contemporary Research in Computer Science and Technology,
2(3), pp. 659-661.