Misplaced Pages

Dark data

Article snapshot taken from Wikipedia with creative commons attribution-sharealike license. Give it a read and then ask your questions in the chat. We can research this topic together.
Data missing or collected but not analysed

Dark data is data which is acquired through various computer network operations but not used in any manner to derive insights or for decision making. The ability of an organisation to collect data can exceed the throughput at which it can analyse the data. In some cases the organisation may not even be aware that the data is being collected. IBM estimate that roughly 90 percent of data generated by sensors and analog-to-digital conversions never get used.

In an industrial context, dark data can include information gathered by sensors and telematics.

Organizations retain dark data for a multitude of reasons, and it is estimated that most companies are only analyzing 1% of their data. Often it is stored for regulatory compliance and record keeping. Some organizations believe that dark data could be useful to them in the future, once they have acquired better analytic and business intelligence technology to process the information. Because storage is inexpensive, storing data is easy. However, storing and securing the data usually entails greater expenses (or even risk) than the potential return profit.

In academic discourse, the term dark data was essentially coined by Bryan P. Heidorn. He uses it to describe research data, especially from the long tail of science (the many, small research projects), which are not or no longer available for research because they disappear in a drawer without adequate data management. Without this, the data become dark, and further reasons for this are e.g. missing metadata annotation, missing data management plans and data curators.

Analysis

The term "dark data" very often refers to data that is not amenable to computer processing. For example, a company might have a great deal of data that exists only as scanned page-images. Even the bare text in such documents is not available without something like Optical character recognition, which can vary greatly in accuracy. Even with OCR, the significance of each part of the data is unavailable. An obvious examples is whether a capitalized word is a name or not, and if so, whether it represents a person, place, organization, or even a work of art. Bibliographic and other references, data within tables (that may be labeled quite adequately for humans, but not for processing), and countless assertions represented with the full complexity and ambiguity of human language.

A lot of unused data is very valuable, and would be used if it could be; but is blocked because it is in formats that are difficult to process, categorise, identify, and analyse. Often the reason that business does not use their dark data is because of the amount of resources it would take and the difficulty of having that data analysed. In other words, the data is "dark" not because it is not used, but because it cannot (feasibly or affordably) be used, given its poor representation.

There are many data representations that can make data much more accessible for automation. However, a great deal of information lacks any such identification of information items or relationships; and much more loses it during "downhill" conversion such as saving to page-oriented representations, printing, scanning, or faxing. The journey back "uphill" can be costly.

According to Computer Weekly, 60% of organisations believe that their own business intelligence reporting capability is "inadequate" and 65% say that they have "somewhat disorganised content management approaches".

Relevance

Useful data may become dark data after it becomes irrelevant, as it is not processed fast enough. This is called "perishable insights" in "live flowing data". For example, if the geolocation of a customer is known to a business, the business can make offer based on the location, however if this data is not processed immediately, it may be irrelevant in the future. According to IBM, about 60 percent of data loses its value immediately.

Storage

According to the New York Times, 90% of energy used by data centres is wasted. If data was not stored, energy costs could be saved. Furthermore, there are costs associated with the underutilisation of information and thus missed opportunities. According to Datamation, "the storage environments of EMEA organizations consist of 54 percent dark data, 32 percent redundant, obsolete and trivial data and 14 percent business-critical data. By 2020, this can add up to $891 billion in storage and management costs that can otherwise be avoided."

The continuous storage of dark data can put an organisation at risk, especially if this data is sensitive. In the case of a breach, this can result in serious repercussions. These can be financial, legal and can seriously hurt an organisation's reputation. For example, a breach of private records of customers could result in the stealing of sensitive information, which could result in identity theft. Another example could be the breach of the company's own sensitive information, for example relating to research and development. These risks can be mitigated by assessing and auditing whether this data is useful to the organisation, employing strong encryption and security and finally, if it is determined to be discarded, then it should be discarded in a way that it becomes unretrievable.

Future

It is generally considered that as more advanced computing systems for analysis of data are built, the higher the value of dark data will be. It has been noted that "data and analytics will be the foundation of the modern industrial revolution". Of course, this includes data that is currently considered "dark data" since there are not enough resources to process it. All this data that is being collected can be used in the future to bring maximum productivity and an ability for organisations to meet consumers' demand. Technology advancements are helping to leverage this dark data affordably. Furthermore, many organisations do not realise the value of dark data right now, for example in healthcare and education organisations deal with large amounts of data that could create a significant "potential to service students and patients in the manner in which the consumer and financial services pursue their target population".

References

  1. ^ "Dark Data". Gartner.
  2. Tittel, Ed (24 September 2014). "The Dangers of Dark Data and How to Minimize Your Exposure". CIO. Archived from the original on 15 January 2019. Retrieved 15 September 2015.
  3. ^ Brantley, Bill (2015-06-17). "The API Briefing: the Challenge of Government's Dark Data". Digitalgov.gov.
  4. ^ Johnson, Heather (2015-10-30). "Digging up dark data: What puts IBM at the forefront of insight economy". SiliconANGLE. Retrieved 2015-11-03.
  5. ^ Dennies, Paul (February 19, 2015). "TeradataVoice: Factories Of The Future: The Value Of Dark Data". Forbes. Archived from the original on 2015-02-22.
  6. Shahzad, M. Ahmad (January 3, 2017). "The big data challenge of transformation for the manufacturing industry". IBM Big Data & Analytics Hub.
  7. "Are you using your dark data effectively". Archived from the original on 2017-01-16. Retrieved 2017-01-12.
  8. Heidorn, P. Bryan. "Shedding light on the dark data in the long tail of science." Library trends 57.2 (2008): 280-299.
  9. Schembera, B., Durán, J.M. Dark Data as the New Challenge for Big Data Science and the Introduction of the Scientific Data Officer. Philos. Technol. 33, 93–115 (2020). https://doi.org/10.1007/s13347-019-00346-x
  10. Miles, Doug (27 December 2013). "Dark data could halt big data's path to success". ComputerWeekly. Retrieved 2015-11-03.
  11. Glanz, James (2012-09-22). "Data Centers Waste Vast Amounts of Energy, Belying Industry Image". The New York Times. Retrieved 2015-11-02.
  12. Hernandez, Pedro (October 30, 2015). "Enterprises are Hoarding 'Dark' Data: Veritas". Datamation. Retrieved 2015-11-04.
  13. "DarkShield Uses Machine Learning to Find and Mask PII". IRI. Retrieved 2019-01-14.
  14. Tittel, Ed (2014-09-24). "The Dangers of Dark Data and How to Minimize Your Exposure". CIO. Archived from the original on 2019-01-15. Retrieved 2015-11-02.
  15. Prag, Crystal (2014-09-30). "Leveraging Dark Data: Q&A with Melissa McCormack". The Machine Learning Times. Retrieved 2015-11-04.
Statistics
Descriptive statistics
Continuous data
Center
Dispersion
Shape
Count data
Summary tables
Dependence
Graphics
Data collection
Study design
Survey methodology
Controlled experiments
Adaptive designs
Observational studies
Statistical inference
Statistical theory
Frequentist inference
Point estimation
Interval estimation
Testing hypotheses
Parametric tests
Specific tests
Goodness of fit
Rank statistics
Bayesian inference
Correlation
Regression analysis
Linear regression
Non-standard predictors
Generalized linear model
Partition of variance
Categorical / Multivariate / Time-series / Survival analysis
Categorical
Multivariate
Time-series
General
Specific tests
Time domain
Frequency domain
Survival
Survival function
Hazard function
Test
Applications
Biostatistics
Engineering statistics
Social statistics
Spatial statistics
Categories: