This shows you the differences between two versions of the page.
| Both sides previous revision Previous revision Next revision | Previous revision | ||
|
dido:public:ra:1.2_views:3_taxonomic:4_data_tax:05_lifecycle:start [2021/10/12 18:18] nick |
dido:public:ra:1.2_views:3_taxonomic:4_data_tax:05_lifecycle:start [2022/05/27 19:51] (current) nick grammar |
||
|---|---|---|---|
| Line 1: | Line 1: | ||
| - | ====== 2.3.4.5 Data Lifecycles ====== | + | ====== 2.3.4.5 Data Lifecycle Taxonomy ====== |
| [[dido:public:ra:1.2_views:3_taxonomic:4_data_tax:start| Return to Data Taxonomy]] | [[dido:public:ra:1.2_views:3_taxonomic:4_data_tax:start| Return to Data Taxonomy]] | ||
| + | ===== Overview ===== | ||
| + | [[dido:public:ra:1.2_views:3_taxonomic:4_data_tax:05_lifecycle:start| Return to Top]] | ||
| - | * [[dido:public:ra:xapend:xapend.a_glossary:d:data_management_platform]] | + | The **Data Lifecycle** covers the stages a particular piece of data transitions through from its initial generation or capture to its eventual archival and/or deletion at the end of its useful life. [[dido:public:ra:xapend:xapend.a_glossary:d:data_management]] is the coordination and administration of all data associated with a project, program or effort. Data Management includes the metadata as well as the individual pieces of data as it progresses through the Data Lifecycle. Data management follows a [[dido:public:ra:xapend:xapend.a_glossary:d:data_strategy]], which includes the required business rules, especially those captured in any governing Legal Documents such as the Charter, By-Laws, and policy and Procedures. Figure {{ref>dataLifecycle}} reflects the major stages in the Data Lifecycle: |
| - | + | - **Create** - The data is created, captured, and copied from other sources | |
| - | The **Data Lifecycle** covers the stages a particular piece of data transitions through from its initial generation or capture to its eventual archival and/or deletion at the end of its useful life. The [[dido:public:ra:xapend:xapend.a_glossary:d:data_management]] is the coordination and administratoin of all data associated with a project, program or effort. DataManagement includes the meta data as well as the inidivaul pieces of data as it progresses through the Data Lifecycle. Data management follows a [[dido:public:ra:xapend:xapend.a_glossary:d:data_strategy]] which also includes the required business rules, especially those captured in any governing Legal Documents such as the Charter, By-Laws, and policy and Procedures. Figure {{ref>dataLifecycle}} reflects the major stages in the Data Lifecycle: | + | - **Store** - The data is stored into a [[dido:public:ra:xapend:xapend.a_glossary:d:datastore]] (i.e., [[dido:public:ra:xapend:xapend.a_glossary:d:database]], [[dido:public:ra:xapend:xapend.a_glossary:d:dom | Document]], Files, etc.) |
| - | + | - **Use** - The Data is actually used, accessed and referenced by an ongoing business process. Often a [[dido:public:ra:xapend:xapend.a_glossary:d:data_management_platform]] is a non-datastore specific way to access data. | |
| - | - **Create** - The data is created, captured, copied from other sources | + | |
| - | - **Store** - The data is stored into a [[dido:public:ra:xapend:xapend.a_glossary:d:datastore]] (i.e., [[dido:public:ra:xapend:xapend.a_glossary:d:database]], [[dido:public:ra:xapend:xapend.a_glossary:d:dom | Document]], Files, etc) | + | |
| - | - **Use** - The Data is actually used, accesed and referenced by an on-going business process | + | |
| - **Propagate** - The Data is copied to other data structures or to other nodes within the system. For example, servers, clients, tiers, and alternative formats (i.e., RDBMS to XML or JSON) | - **Propagate** - The Data is copied to other data structures or to other nodes within the system. For example, servers, clients, tiers, and alternative formats (i.e., RDBMS to XML or JSON) | ||
| - **Share** - The data is made public and can be shared with other internal or external business processes using any number of mechanisms such as <WRAP> | - **Share** - The data is made public and can be shared with other internal or external business processes using any number of mechanisms such as <WRAP> | ||
| [[dido:public:ra:xapend:xapend.a_glossary:h:https|HTTP(S)]], | [[dido:public:ra:xapend:xapend.a_glossary:h:https|HTTP(S)]], | ||
| - | [[dido:public:ra:xapend:xapend.a_glossary:f:ftp]], SMTP, SMS, IPFS, DDS, | + | [[dido:public:ra:xapend:xapend.a_glossary:f:ftp]], SMTP, [[dido:public:ra:xapend:xapend.a_glossary:s:sms]], [[dido:public:ra:xapend:xapend.a_glossary:i:ipfs]], [[dido:public:ra:xapend:xapend.a_glossary:d:dds]], |
| ]]dido:public:ra:xapend:xapend.a_glossary:b:blockchain]], | ]]dido:public:ra:xapend:xapend.a_glossary:b:blockchain]], | ||
| [[dido:public:ra:xapend:xapend.a_glossary:d:dlt]], etc.</WRAP> | [[dido:public:ra:xapend:xapend.a_glossary:d:dlt]], etc.</WRAP> | ||
| - **Archive** - The data is preserved as part of the system legacy. In the past, the amount of data archived was limited primarily due to the cost of storage, but with the advent of inexpensive offline storage, most data is now preserved. | - **Archive** - The data is preserved as part of the system legacy. In the past, the amount of data archived was limited primarily due to the cost of storage, but with the advent of inexpensive offline storage, most data is now preserved. | ||
| - | - **Destroy** - The data is considered as having no value, become a laibility or is required to be destroyed by law (i.e., The Right to Be Forgotten). In traditional data systems, the data is usually overwritten with newer, more germain data. For example, old room temperatures are replaces with new ones if there is no requirement to archive the old temperature. | + | - **Destroy** - The data is considered as having no value, has become a liability or is required to be destroyed by law (i.e., The Right to Be Forgotten). In traditional data systems, the data is usually overwritten with newer, more germane data. For example, old room temperatures are replaced with new ones if there is no requirement to archive the old temperature. |
| - | : **Note:** Figure {{ref>dataLifecycle}} also includes a **Plan** Stage, however, since the data does not exist during planning, it is not considered as an actual Data Sage, but this in no way means it is not an important stage for data. It is during this stage the requirements for the data are identfied using system architecture and engineering and modeling. Planning usually includes [[dido:public:ra:xapend:xapend.a_glossary:u:use_case | Use-Cases]], [[dido:public:ra:xapend:xapend.a_glossary:p:prototype | Prototypes]], and the various [[dido:public:ra:1.2_views:3_taxonomic:4_data_tax:04_modeltypes | Data Models]] (i.e., Conceptual, Logical and Physical), [[dido:public:ra:1.3_gov:1_legaldocs | legal conciderations]] and the pragmatics of things like the quanity and quality of data. Sometimes during this stage, data that is "planned" never comes to fruition. | + | : **Note:** Figure {{ref>dataLifecycle}} also includes a **Plan** Stage, however, since the data does not exist during planning, it is not considered as an actual Data Stage and, in no way means, it is not an important stage for data. It is during this stage the requirements for the data are identified using system architecture and engineering and modeling. Planning typically includes [[dido:public:ra:xapend:xapend.a_glossary:u:use_case | Use-Cases]], [[dido:public:ra:xapend:xapend.a_glossary:p:prototype | Prototypes]], and the various [[dido:public:ra:1.2_views:3_taxonomic:4_data_tax:04_modeltypes:start| Data Models]] (i.e., Conceptual, Logical and Physical), [[dido:public:ra:1.3_gov:1_legaldocs | legal considerations]] and the pragmatics of things like the quality and quality of data. Sometimes during this stage, the "planned" data never comes to fruition. |
| <figure dataLifecycle> | <figure dataLifecycle> | ||
| Line 27: | Line 26: | ||
| </figure> | </figure> | ||
| + | Wigmore(( | ||
| + | Ivy Wigmore, | ||
| + | __data Life Cycle__, | ||
| + | TechTarget, | ||
| + | July 2017, | ||
| + | Accessed: 12 October 2021, | ||
| + | [[https://whatis.techtarget.com/definition/data-life-cycle]] | ||
| + | )) defines only six stages for the Data Lifecycle: | ||
| + | - **Generation** or **capture**: In this phase, data comes into an organization, usually through data entry, acquisition from an external source, or signal reception, such as transmitted sensor data. | ||
| + | - **Maintenance**: In this phase, data is processed prior to its use. The data may be subjected to processes such as integration, scrubbing, and extract-transform-load (ETL). | ||
| + | - **Active use**: In this phase, data is used to support the organization’s objectives and operations. | ||
| + | - **Publication**: In this phase, data isn’t necessarily made available to the broader public but is just sent outside the organization. Publication may or may not be part of the life cycle for a particular unit of data. | ||
| + | - **Archiving**: In this phase, data is removed from all active production environments. It is no longer processed, used, or published but is stored in case it is needed again in the future. | ||
| + | - **Purging:** In this phase, every copy of data is deleted. Typically, this is performed on data that is already archived. | ||
| - | ----- | + | : **Note:** These six stages are roughly the same as the ones in Figure {{ref>dataLifecycle}} but with slightly different names. Wigmore's model combines the **Propagation** and **Sharing** stages into a single **Publication** stage. |
| + | : **Note:** It is important to have a [[dido:public:ra:xapend:xapend.a_glossary:d:data_retention_policy]] in place during all the stages of the Data Life Cycle. During **Planning**, knowing the Data Retention Policies governing the data can influence the kinds of information collected and how the data is stored. For example, partitioning the data by it's data retention policy allows data with similar governing policies to be **Archived** and **Deleted** at the same time while retaining other less sensitive data. | ||
| - | Although specifics vary, data management experts often identify six or more stages in the data life cycle. Here's one example: | + | ===== DIDO Specifics ===== |
| - | + | [[dido:public:ra:1.2_views:3_taxonomic:4_data_tax:05_lifecycle:start| Return to Top]] | |
| - | Generation or capture: In this phase, data comes into an organization, usually through data entry, acquisition from an external source or signal reception, such as transmitted sensor data. | + | |
| - | Maintenance: In this phase, data is processed prior to its use. The data may be subjected to processes such as integration, scrubbing and extract-transform-load (ETL). | + | |
| - | Active use: In this phase, data is used to support the organization’s objectives and operations. | + | |
| - | Publication: In this phase, data isn’t necessarily made available to the broader public but is just sent outside the organization. Publication may or may not be part of the life cycle for a particular unit of data. | + | |
| - | Archiving: In this phase, data is removed from all active production environments. It is no longer processed, used or published but is stored in case it is needed again in the future. | + | |
| - | Purging: In this phase, every copy of data is deleted. Typically, this is performed on data that is already archived. | + | |
| - | Data lifecycle management (DLM) is becoming increasingly important since the explosion of big data and the ongoing development of the Internet of Things (IoT). Enormous volumes of data are being generated by an ever-increasing number of devices all over the world. Proper oversight of data throughout its life cycle is essential to optimize its usefulness and minimize the potential for errors. Finally, archiving or deleting data at the end of its useful life ensures that it does not consume more resources than necessary. | + | |
| + | The main differences betwen a generic Data Lifecycle and a DIDO Data Lifecycle is that in the idealized DIDO Data Lifecycle the **Destroy** and **Archive** stages are non-existent or modified. See Figure {{ref>didoLidecycle}}. | ||
| <figure didoLidecycle> | <figure didoLidecycle> | ||
| Line 45: | Line 53: | ||
| <caption>Immutable Data Lifecycle.</caption> | <caption>Immutable Data Lifecycle.</caption> | ||
| </figure> | </figure> | ||
| + | |||
| + | However, because the data within a DIDO is theoretically immutable and no data is ever lost, the panacea of "unlimited" data storage is being challenged. There are different ways that the size of the ledgers can be managed. One way is to have different kinds of nodes with each node containing differing amounts of data. See section [[dido:public:ra:1.2_views:3_taxonomic:3_node_tax:start]]. | ||
| + | |||
| + | <figure nodeTaxonomy> | ||
| + | {{ dido:public:ra:1.2_views:1_stake:3_taxonomic:node_taxonomy.png?550 |}} | ||
| + | <caption>DIDO Node Taxonomy: Node Types | ||
| + | </caption> | ||
| + | </figure> | ||
| + | |||
| + | |||
| + | Another mechanism to overcome the "size" issue is to use [[dido:public:ra:xapend:xapend.a_glossary:s:sharding]]. Sharding splits a database (in a DIDO the Ledger) horizontally into [[dido:public:ra:xapend:xapend.a_glossary:s:shard | Shards]] spreading the load across the many Shards. In an Ethereum context, Sharding is intended to reduce network congestion and increase transactions per second by creating new Shards referred to as **Chains**. | ||
| + | |||
| + | Iota is using a form of Sharding called Streams and has an [[dido:public:ra:xapend:xapend.a_glossary:r:rfp]] at the [[dido:public:ra:xapend:xapend.a_glossary:o:omg]] to provide a standardized specification for publishing, subscribing and validating the [[dido:public:ra:1.2_views:2_tech_views:2-nodenet:3_nodearch:2_ido:2_trans | transactions]] on the "Tangle" (Iota's version of [[dido:public:ra:xapend:xapend.a_glossary:d:directed_acyclic_graph_dag]]). (( | ||
| + | Object Management Group (OMG), | ||
| + | __Linked Encrypted Transaction Streams (LETS)__, | ||
| + | Accessed: 17 October 2021, | ||
| + | [[https://www.omg.org/hot-topics/linked-encrypted-transaction-streams.htm]], and the RFP is here: | ||
| + | [[https://www.omg.org/cgi-bin/doc.cgi?mars/20-12-22]] | ||
| + | )), | ||
| + | (( | ||
| + | Perry Cohen, | ||
| + | Embedding Computing, | ||
| + | __OMG Issues RFP to Standardize Linked Encrypted Transaction Streams (LETS)__, | ||
| + | January 2021, | ||
| + | Accessed: 17 October 2021, | ||
| + | [[https://www.embeddedcomputing.com/technology/security/network-security/omg-issues-rfp-to-standardize-linked-encrypted-transaction-streams-lets]] | ||
| + | )) | ||
| + | |||
| + | <color blue><todo @char #char:2022-03-20>New Section -- review </todo></color> | ||
| /**=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=- | /**=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=- | ||