A Category in the Learn & Work Ecosystem Library

Privacy & Data Protection 44

Show All Definitions

An artificial intelligence system (bot or AI avatar) deployed to attend a meeting or event in place of a human, typically to observe, record, summarize, or interact on their behalf. The “emissary” may be pre-programmed with instructions or agenda points, or be a real-time agent capable of basic interactions or just recording. The AI may appear as a name or video avatar in a Zoom/Teams screen and may or may not clearly indicate it is not human. The AI listens or records the meeting, may generate a summary or transcript afterward, and may or may not participate (e.g., say “noted,” “I’ll follow up.”  The human host of the meeting can typically recognize this is an AI representative but may or may not disclose this substitution to all other participants.

This is an emerging and controversial practice in AI-mediated communication that currently lacks a standardized term—but several terms are circulating in technology, business, and educational communities to describe the practice:

  • AI Representative | AI Emissary
    • Not standard terms
    • AI Representative implies delegation (like a spokesperson)
    • AI Emissary has a formal, more diplomatic tone
    • Both suggest the AI is attending on behalf of a human, not pretending to be one
  • Digital Double
    • Popular in speculative tech and AI circles
    • Suggests a full virtual representation of a person
    • Can imply real-time interaction or decision-making capability
    • Closer to Digital Twin or Avatar with Agency
  • Ghost Attendee or Proxy Bot
    • Colloquial terms, sometimes used pejoratively
    • Emphasizes the ethical ambiguity—present but not present
    • Proxy bot may refer to a simple note-taking tool
  • Shadow Presence
    • Describes situations where a bot is present in a digital meeting space (or even listens via device microphone), but its presence is not openly acknowledged
    • Often unintentional but raises ethical flags
  • Synthetic Identity
    • An AI-generated name, avatar, or persona used to represent a bot in professional settings
    • Can be used transparently (e.g., “Fireflies.ai bot”) or misleadingly (e.g., a human-sounding name with no disclosure it’s AI)
  • Meeting Clone or AI Meeting Attendee
    • Used in some productivity/startup platforms
    • Describes tools like Otter.ai, Rewind, Fathom, and Zoom AI Companion that show up with your name and join calls

Citation Note: This glossary entry was developed by the Learn & Work Ecosystem Library with support from ChatGPT (OpenAI, 2025) and is based on an original synthesis of emerging terminology from journalism, technical documentation, and academic discourse. The terms reflect evolving usage in the fields of AI, workplace automation, and digital communication, and are not yet standardized across sectors.

In badging, refers to representations that share information about a badge belonging to one earner. Assertions are packaged for transmission as JSON objects with a set of mandatory and optional properties. The assertion typically includes information about who earned the badge, what the badge represents, and who issued the badge. The assertion for a badge includes data items required by the Open Badges Specification: unique ID, recipient, badge URL, verification data, issue date. Optional data items include a badge image with assertion data baked into it, evidence URL, expire date, and stored location in a hosted file or JSON Web signature.

According to the UN Refugee Agency, when people flee their own country and seek sanctuary in another country, they apply for asylum. This is the right to be recognized as a refugee and receive legal protection and material assistance. An asylum seeker must demonstrate that their fear of persecution in their home country is well-founded.

Related terms: Refugee, internally displaced person, stateless person

The process of embedding assertion data into a badge image.

Closed-circuit AI models are artificial intelligence (AI) systems designed to operate within restricted, organization-controlled environments. They are trained or configured using limited, curated datasets specific to a defined domain, organization, or set of tasks. These models are intentionally constrained in scope, data access, and behavior to meet requirements related to privacy, compliance, reliability, and operational control, especially by companies that commission closed-circuit models.

Employers often adopt closed-circuit AI models for use cases in human resources, legal review, training systems, and internal knowledge management, particularly where sensitive data, regulatory obligations, or reputational risk are involved. While these models typically offer less breadth than general large language models (LLMs), they provide greater governance, predictability, and alignment with organizational requirements.

Closed-circuit does not mean that the information received from AI is more accurate or transparent; rather, it reflects a design choice that prioritizes control and accountability over open-ended exploration. Many organizations increasingly use hybrid approaches, combining the breadth of LLMs with closed-circuit controls. Employers in regulated and high-risk environments (e.g., healthcare, finance, law) increasingly favor closed-circuit AI deployments due to data sensitivity, compliance requirements, and reputational considerations. Consequently, many organizations limit general-purpose AI tools to exploratory or low-risk tasks, reserving operational decision-making for constrained systems.

See Glossary: Large Language Models (LLMs) | Learn & Work Ecosystem Library

See Topic Brief: AI Architectures in the Workplace: Large Language Models & Closed-Circuit AI Systems | Learn & Work Ecosystem Library

As defined by the Data Quality Campaign in Data 101, cross-agency data governance is a formal, leadership-level body that is responsible and accountable for making decisions about how data linked between state agencies is connected, secured, accessed, and used to meet state education and workforce goals. Typical features of these entities:

  • They define clear purposes, roles, and responsibilities for participating agencies in the statewide longitudinal data system (SLDS) and ensure accountability for data quality, privacy, and security.
  • The entity is codified into state law to ensure the right membership and sustainability over time,
  • Best-practice includes senior leaders from each of the agencies that contribute data to the system plus stakeholders that represent state education and workforce priorities (e.g., business leaders, district leaders, community organizations).
  • State leaders set a vision for data use and create accountability for making decisions about data.
  • Creates sustainability for the SLDS, especially when codified by legislation, by ensuring that decision-making authority is clear and responsibilities for data collection, privacy and security, and access are defined.
  • Builds trust by creating a space for cross-agency collaboration, facilitating a shared data culture, and ensuring processes and decisions are transparent.
  • Creates forums for communication and decision-making that are open to and include input from the public as well as local government.

As defined by IBM, a data center is a physical room, building or facility that houses IT infrastructure for building, running and delivering applications and services. It also stores and manages the data associated with those applications and services.

Historically, data centers were privately owned, tightly controlled on-premises facilities housing traditional IT infrastructure for the exclusive use of one company. Recently, they have evolved into remote facilities or networks of facilities owned by cloud service providers (CSPs). These CSP data centers house virtualized IT infrastructure for the shared use of multiple companies and customers.

Managed data centers and colocation facilities are options for organizations that lack the space, staff or expertise to manage their IT infrastructure on-premises, and used by those who prefer not to host their infrastructure by using the shared resources of a public cloud data center. Companies often choose managed data centers and colocation facilities to house remote data backup and disaster recovery (DR) technology for small and mid-sized businesses (SMBs).

  • In a managed data center, the client organization leases dedicated servers, storage and networking hardware from the provider. The provider handles the client’s administration, monitoring and management.
  • In a colocation facility, the client owns all the infrastructure and leases a dedicated space to host it within the facility. In the traditional colocation model, the client organization has sole access to the hardware and full responsibility for managing it. This model is ideal for privacy and security but can be impractical, particularly during outages or emergencies. Today, most colocation providers offer management and monitoring services to clients who want them.

There are different types of data center facilities:

  • Enterprise (on-premises) data centers— The user organization is responsible for all deployment, monitoring and management tasks. These centers offer users more control over information security and can more easily comply with regulations such as the European Union General Data Protection Regulation (GDPR)or the US Health Insurance Portability and Accountability Act (HIPAA).
  • Public cloud data centers and hyperscale data centers—House IT infrastructure resources for shared use by multiple customers (from several to millions) through an internet connection. Many of the largest cloud data centers (known as hyperscale data centers) are run by major cloud service providers (CSPs), such as Amazon Web Services (AWS), Google Cloud Platform, IBM Cloud, and Microsoft Azure. These companies have major data centers in every region of the world. Hyperscale centers are larger than traditional data centers and can cover millions of square feet. They typically contain at least 5,000 servers and miles of connection equipment, and can be as large as 60,000 square feet.
  • Edge data centers (EDCs) —Cloud service providers typically maintain smaller data centers that are located closer to cloud customers and their customers as well. These centers form the foundation for edge computing, a distributed computing framework that brings applications closer to end users. These centers are useful for real-time, data-intensive workloads like big data analytics, AI, machine learning, and content delivery.

A study from McKinsey & Company projects the industry to grow at 10% a year through 2030, with global spending on the construction of new facilities reaching USD49 billion.

According to the Data Quality Campaign’s Data 101, there are three key types of data common to education and workforce:

  • Student Data (P-12)
    • Information such as attendance, grades, student growth, outcomes, enrollment.
    • Schools, school districts, and states collect student data and use it to make decisions about instruction, interventions, policy development, and resource allocation.
    • Most student data is stored at the school and district levels.
    • A limited amount of this data is reported to states.
    • States use and share this data in anonymous and aggregate forms.
  • Postsecondary Data
    • Information such as admissions, enrollment, persistence, completion, courses or majors, use of support services or public benefits, financial aid, and student debt.
    • Information is collected to help better understand education options after high school, such as two- and four-year college programs, applied career training through community and technical colleges, professional certifications, and other noncredit training or education pursued after high school.
    • Many state, federal, and accrediting agencies require postsecondary institutions to report data for compliance purposes (e.g., federal financial aid eligibility or state authorization requirements).
    • Reporting requirements by governmental and other regulatory entities shape much of the data collected in the field.
    • Higher education institutions collect additional data on individual students and education programs. Data are typically used to evaluate programs and improve the tools and support offered to students. Such data is rarely collected at the state or federal levels.
  • Workforce Data
    • Information relating to local or regional labor markets such as data about the existing workforce; wage information; use of social assistance programs; demand for occupations; growing fields and industry trends; and information on available training, skills, and credentialing resources.
    • Comes from a variety of sources, such as state unemployment records, tax data, public benefits programs, job posting sites, education institutions, job training programs, apprenticeship programs, the military, and adult education services.
    • The amount of data incorporated into statewide data systems is often limited. For example, in states where the data is incorporated, it may include only unemployment insurance data, which contains information from employers on their employees and their wages but has many gaps and limitations.
    • In many states, much of the data may be available only to private entities and employers.

Encryption is a security process to convert data into an unreadable code. The purpose is to ensure that sensitive data is protected from unauthorized access. Encryption happens when data is being transmitted or stored to keep it secure. The process uses complex mathematical algorithms, to convert readable data (plaintext) into unreadable data (ciphertext). Only those with the decryption key is able to reverse the encryption process to access the original data. Encryption is widely used in online transactions, communication, and data storage, to protect against unauthorized access and cyber threats.

Decryption is the reverse process. It converts encrypted data back to its readable form (plaintext) using the correct decryption key. Decryption takes place when the intended recipient receives the encrypted data and uses the proper key to restore the original message.

These two processes work together to safeguard data from unauthorized access and tampering.

According to Digital Promise, data governance refers to a data management process or framework whereby an organization or a collection of organizations ensure the collection, storage, and security of high-quality data.

Data literacy and data science skills are important for jobs in fields in business, engineering mathematics, statistics, computer science, life sciences, social sciences, digital humanities, and others. They are also important skills for navigating an increasingly data-driven world. These terms are useful especially to education and training programs established to develop learners’ skills, K-12 through postsecondary education.

Data literacy refers to foundational knowledge about and the ability to read, understand, and communicate data or claims derived from data. Literacy includes knowing how to communicate and question data and representations of data critically, including limitations and potential biases. Areas of knowledge include understanding probability and randomness, ways to visualize data, the meaning of descriptive statistics, the concept of statistical significance, the concept of mathematical techniques called hypothesis tests, how data is collected, and the concept of control groups.

Data science is an interdisciplinary field that refers to applying the processes of working with data. Applications can include calculating means and medians, formatting and graphing data including with multiple variables, performing a scientific study which includes establishing framing questions and crafting the methods of data collection, and measuring variation among the data collected.

According to Digital Promise, data sharing systems are technological platforms where data collected by multiple entities can be shared across and within multiple stakeholders’ organizations.

Refers to having confidence that data meets quality standards and is ready to act on (e.g., analyze, understand, make decisions about). Data trust cannot be taken for granted; there should be evidence of the data’s quality standards.

The Data Management Association of the UK defines six dimensions of data quality required for trustworthy data:

  1. Accuracy —degree to which data correctly describes the real-world object or event in question.
  2. Completeness —proportion (percentage) of data stored versus 100% complete (e.g., are there blank values indicating certain data has not been populated).
  3. Consistency —absence of difference when comparing two or more representations of an item against a definition (e.g., do dates and names match).
  4. Timeliness — extent to which data is current enough to represent reality for realistic use.
  5. Uniqueness — no item or entity instance is recorded more than once based upon how that item is identified (nonduplication).
  6. Validity or conformity — extent to which data conforms to the syntax (format, type, range) of its definition.

There is important context behind data trust: (1) the multiplicity of data sources, and (2) potential errors in data management:

  • Manual data quality management cannot handle immense volumes of data that come from multiple sources of data; e.g., from SaaS and web applications, direct data entry like web forms, unstructured sources like social media posts, machines such as smartphones and “Internet of Things” devices.
  • Errors in data management are a factor in quality data. Errors are introduced by people who make mistakes, machines that are not foolproof, and errors can be introduced by data that often passes through complex information systems coded by multiple developers.

Refers to a social web in which no single entity (or small group of entities) controls the Internet. The decentralized web is focused on principles of self-determination of users who have their own data and join the network seeking new practices around sovereignty, openness, protection of their data, and censorship-resistance. The decentralized Internet is also known as Web3 or the Dweb. Unlike the current Web (Web2) which relies on centralized servers, clouds, and platforms, Web3 embraces blockchain, peer-to-peer networks, and distributed storage. In Web3, ownership and control are distributed among users. This places ownership of personal data back into the hands of individuals rather than companies, which many contend take data from individuals and track both user’s data and data from user’s networks without permission.

A secure mobile app that digitally stores and manages personal information and identity. These credentials can include IDs, passports, driver’s licenses, and other verifiable forms of identification. Unlike traditional identity systems that depend on centralized parties and external databases, digital ID wallets can store data directly on the user’s device, enabling individuals to determine what information to share, with whom, and under what conditions. This can ensure a higher level of privacy and security, as it minimizes data exposure and reduces reliance on third-party data storage.

Digital identity wallets eliminate the need for physical documents, streamlining identity verification and management. Governments worldwide are increasingly adopting digital ID solutions to enhance citizen interactions, improve service efficiency, and reduce fraud.

See Topic: Growth of Digital Tools: Digital Wallets, Skills Passports, and Digital ID Wallets | Learn & Work Ecosystem Library

Digital provenance (sometimes called data provenance or data lineage when referring specifically to datasets) is the documented history of a digital asset or data record—how it was created, modified, shared, and used over time. Provenance records help people verify whether digital content or data is authentic, reliable, and trustworthy by showing who created it, when changes occurred, and whether it has been altered.

Provenance can be captured through secure metadata, cryptographic signatures, audit logs, or—in some cases—blockchain systems that create tamper-resistant records. While blockchain can strengthen provenance, it is not required; the core purpose of digital provenance is to make the origins and lifecycle of digital information transparent for verification and trust.

In the learn-and-work ecosystem, provenance is particularly important for digital credentials, Learning and Employment Records (LERs), workforce data, and research datasets. It allows educational institutions, employers, and learners to understand and verify the origin, lifecycle, and integrity of information.

Provenance and verification are related terms: provenance is the historical evidence that shows an asset’s origin and transformations; verification is the process of checking that evidence to confirm authenticity.

See Topic Brief: Digital Provenance | Learn & Work Ecosystem Library

Allows data to be processed and analyzed closer to the source of the data, rather than in a centralized data center. This can improve response times, reduce latency (amount of time it takes for a data packet to travel from one designated point to another), and reduce the amount of data that needs to be transferred over connected networks.

According to the Data Quality Campaign, education data is information about individuals, groups, and entire populations. This includes:

  • course access and attendance by students
  • performance data and postsecondary enrollment rates
  • any information that can be used to support individuals throughout their education and workforce journeys.

Refers to the repeated use of digital technologies to monitor, harass, intimidate, threaten, or otherwise target an employee in ways that cause fear, emotional distress, or disruption to their work or personal life.

Cyberstalking may be committed by coworkers, supervisors, former employees, customers, clients, or unrelated individuals and can occur through email, text messages, social media, collaboration platforms, online forums, location-tracking technologies, or other digital communication tools.

Unlike isolated online harassment, cyberstalking involves a persistent pattern of unwanted behavior and may escalate into workplace safety, legal, or security concerns.

Organizations increasingly address employee cyberstalking through workplace violence prevention, anti-harassment, cybersecurity, employee well-being, and acceptable technology use policies, often in coordination with legal actions.

See Topic Brief: Digital Workplace Safety | Learn & Work Ecosystem Library

Fair use is a legal doctrine in U.S. copyright law that permits limited use of copyrighted material without permission from the rights holder for purposes such as education, research, scholarship, commentary, criticism, and news reporting. Fair use determinations are context-specific and rely on an analysis of four factors: (1) the purpose and character of the use, (2) the nature of the copyrighted work, (3) the amount used, and (4) the effect on the market value of the original work.

The growth of artificial intelligence (AI) has complicated traditional understandings of fair use, raising new legal and institutional questions about the use of copyrighted materials.  In digital learning, research, and AI-enabled environments, fair use plays a critical role in enabling access, innovation, and knowledge sharing while balancing the rights of content creators. AI complicates fair use practices because it raises unresolved questions such as:

  • Whether training on copyrighted works constitutes fair use?
  • Whether AI outputs are “transformative”?
  • Who bears responsibility when outputs resemble copyrighted material?
  • How should attribution, licensing, and compensation work at scale?

Courts have historically emphasized human purpose and transformation. AI introduces non-human intermediaries, probabilistic reuse, and scale far beyond traditional educational copying. This matters for educators, researchers, libraries, and publishers because fair use is no longer just about what humans copy: it is about what systems ingest, how outputs are generated, and how institutions manage risk, disclosure, and compliance.

Refers to data that typically includes an individual’s wage and salary information, bank account information, savings information, and credit score.

A student loan in which students receive money to fund their education or training. Students agree via a contract agreement to pay the ISA provider a fixed percentage of their income for a set period of time after they finish school and pass a specific income threshold. They may repay more or less than the amount received, depending on the agreement’s terms. If the student later loses his/her job, the terms typically permit the individual to stop making payments. Although ISA providers often advertise their products as an alternative to loans, the Consumer Financial Protection Bureau (a federal regulatory agency) has found that ISAs are student loans.

The Learn & Work Ecosystem Library functions, in part, as an information aggregator database, as it primarily comprises resources that consist of processed, organized, and structured information. This information is data that has been analyzed, categorized, and linked to other relevant content, making it meaningful and actionable for a variety of stakeholders exploring the learn-and-work ecosystem.

To fully grasp the concept of an information aggregator database, it is essential to understand the key elements: data, databases, aggregation, information, and knowledge:

  • Data refers to raw, unprocessed facts or figures, often numerical or descriptive, that lack inherent meaning by themselves. These are typically collected from observations or measurements.
  • Database is an organized collection of structured data, typically stored electronically. Databases support various functions such as accessing, managing, updating, organizing, linking, and analyzing data. Organizations often maintain multiple databases for distinct purposes, like tracking sales, inventory, or transactions, while some specialize in aggregating data from various sources to support deeper analysis and reporting.
  • Data aggregation is the process of gathering and compiling data from different sources into a single, unified dataset. This is commonly used in business intelligence, analytics, and data science to derive insights, identify trends, and support strategic decision-making. Aggregated data is often summarized to facilitate easier analysis and interpretation.
  • Information is data that has been processed or structured in a way that makes it meaningful and useful. It consists of facts or details about a specific topic or event that can be communicated and understood.
  • Knowledge is the result of interpreting and synthesizing information within a particular context. It represents the human capacity to process and combine information, transforming it into deeper understanding. Individuals express their knowledge by encoding it as information—through books, websites, or other mediums—that others can use to build their own knowledge.

In essence, the Learn & Work Ecosystem Library organizes and curates diverse data and information to create a comprehensive, interconnected resource. This makes it easier for users to navigate, understand, and act on the vast landscape of learning and workforce development.

See Glossary: Web Scraping & Information Aggregators | Learn & Work Ecosystem Library

See Topic: Library as Information Aggregator Database: Continuum Model of Data to Information, Information to Knowledge, Knowledge to Wisdom | Learn & Work Ecosystem Library

According to the UN Refugee Agency, an IDP is a person who has been forced to flee their home but never cross an international border. IDPs include people displaced by internal strife and natural disasters. These individuals seek safety wherever they can—in nearby towns, schools, settlements, internal camps, forests and fields. Unlike refugees, IDPs are not protected by international law or eligible to receive many types of aid because they are legally under the protection of their own government.

Related terms: Refugee, stateless person, asylum seeker

An international classification system developed and maintained by the International Organization for Standardization.  ICS are used to catalog and classify standards,  often for use in databases and libraries.  ICS currently includes 40 fields. Standards are organized according to:

  • sectors of the economy (e.g., agriculture, mining construction, packaging industry)
  • technologies (e.g., telecommunications, food processing)
  • activities (e.g., environmental protection , safety assurance and protection of public health)
  • fields of science (e.g., mathematics, astronomy)

The latest editions of the ICS are downloadable free of charge from the ISO website. Anyone may propose revisions or additions to the ICS.

Refers to a network of physical devices, vehicles, appliances, and other physical objects that use sensors, software, and network connectivity to collect and exchange data over the internet or other communications networks.  IoT devices are often known as “smart objects” because they have interconnection capacity to share data. Such devices typically include sensors and actuators; connectivity technologies; cloud computing; big data analytics; and security and privacy technologies. IoT applications are prevalent in healthcare, manufacturing, retail, agriculture, and transportation.

A knowledge graph is a design pattern that can help users better understand data.  A graph is formed of nodes and relationships:

  • a node is a person, object, location, or event
  • relationship refers to the interaction among nodes.

Representing data to depict these connections (relationships among nodes) can enhance the value of information.

According to Wikipedia, the Knowledge Graph was launched by Google in May 2012 to enhance the value of information returned through Google searches. By May 2016, knowledge boxes appeared for some one-third of the 100 billion monthly searches the company was processing.  The Google Knowledge Graph depicts relevant information in an infobox next to its search results. This enables users to see the answer in a glance. Data is generated automatically from multiple sources, and it covers places, people, organizations, and more.

Knowledge graphs are at the core of human-facing technologies such as search, question answering, dialogue, and recommenders.  They are particularly useful in fields characterized by data silos (e.g., healthcare and financial services). Knowledge graphs can help with data governance, fraud detection, knowledge management, searches, chatbot use, developing recommendations, and developing intelligent systems across different organizational units.

Refer to entry-level positions that are accessible to individuals without a college degree that offer some combination of: (1) higher-than-average starting pay and continued wage premiums 10 years on; (2) higher likelihood of providing health insurance; (3) high promotion rates (transitions from a first job that are a step up in responsibility or pay); (4) strong pathways to better opportunities; (5) some level of protection from technological disruption (low automation risk).

See: Research: “Launchpad Jobs: Achieving Career and Economic Success Without a Degree” — American Student Assistance & Burning Glass Institute | Learn & Work Ecosystem Library

Refers to information that describes, organizes, and connects other information. It is often called “data about data” because it helps explain what something is, how to find it, and how it is structured. Metadata is used in libraries, websites, and digital systems to make information easier to search for and use.

There are different types of metadata, including:

  • Descriptive metadata (like a book’s title, author, and keywords)
  • Access metadata (information that helps people find and use a resource)
  • Structural metadata (details about how information is organized or linked)

Metadata is also known as cataloging information, indexing data, or linked data when it connects different pieces of information in a way both people and computers can understand.

At the Learn & Work Ecosystem Library, all public resource descriptions are metadata.

Multi-state data collaboratives are regional networks (coalitions) of state workforce, education, human services, and other agencies that partner with each other and with regional postsecondary education partners to produce data products that policymakers, practitioners, and citizens use to inform policy developments and answer questions critical to society. An example is the Multi-State Data Collaborative (MSDC) led by the National Association of Workforce Agencies (NASWA).

As defined by CredLens, refers to an entity with a defined legal and technical framework for managing and governing data on behalf of a group or community, with a focus on national or large-scale public interest data. ‍The trust acts as a steward of the data, often with the goal of enhancing transparency, fostering innovation and supporting public good initiatives while ensuring compliance with privacy laws and regulations.

The standard used by Federal statistical agencies in classifying business establishments for the purpose of collecting, analyzing, and publishing statistical data related to the U.S. business economy.

Open Data is data that can be freely used and distributed by anyone to use, reuse, and redistribute. It requires data to be available, which means the data must be in the public domain or under license conditions that allow users to use the data without restriction.

Linked Data are datasets that make use of clear, unique identifiers that allow elements and relationships in different datasets to be identified as referring to the same thing. Linked data systems are built on a standard of World Wide Web (Web or Internet) technologies which enable linked data to become a global database. Linked Data enables links between datasets that are understandable to both humans and machines.

Open Data is not the same as Linked Data. Linked Data need not be Open Data, and Open Data need not be Linked Data.  Open Data is available to everyone without links to other data. At the same time, data can be linked without being freely available for reuse and distribution. Linked Data is enabled through a set of design principles for sharing machine-readable interlinked data on the Web. Linked Data enables the creation of a global network of data. The resulting network can automatically answer complex queries and analytics by searching the network of information and finding matches (known as Data Graph Traversal).

Linked Open Data is a blend of Linked Data and Open Data: it is both linked and uses open sources. The benefits of Linked Open Data: it breaks down information silos between various formats (often disparate sources and formats), facilitates the extension of data models, and allows for easy updates. The resulting data integration enables easier and more efficient searching through complex databases for information.

Main principles of the Linked Open Data:

  • Uniform Resource Identifiers (URIs) are used to name and identify individual things. The URI is a single global identification system used to give unique names to anything (e.g., resources on a webpage; mail address; phone number; books; objects such as people and places; and concepts).
  • All conceptual things have a name starting with HTTP (Hypertext Transfer Protocol)—an application (software) protocol in the Internet’s model of distributed, collaborative, hypermedia information systems. HTTP is the foundation of data communication for the Web, where hypertext documents include hyperlinks to other resources that users can easily access.
  • Looking up an HTTP name returns useful data about the thing in question in a standard format. Anything else that that same thing has a relationship with through its data also has a name beginning with HTTP.
  • When publishing data on the Web, other things are referred to using their HTTP URI-based names.
  • URIs provide a means of locating and retrieving information resources on a network (either on the Internet or private network, such as a computer.

Refers to services that are purchased from external providers, i.e., “what you pay someone else to do.” Examples include externally provided help desk, data center, or services provided by multicampus system or district offices; food services; cleaning services; a range of one-time project costs and professional services.

 

P20W refers to pre-kindergarten through college and into the workforce.

P20W data systems refer to the various state-level educational databases that collect student data across pre-kindergarten through college and into the workforce to help education leaders make policy decisions, the best use of resources, and support individuals throughout their life stages.

State Longitudinal Data Systems (SLDS) connect statewide information from early childhood through K–12 education, postsecondary education, and the workforce. These state-level data infrastructures in the U.S. securely bring together cross-agency data that enable leaders, practitioners, and community members to better understand the progress, predictors, and performance of learners throughout their educational and employment pathways. According to the Data Quality Campaign, SLDS:

  • Reside in different places depending on the state context, but best practice is that the SLDS itself is not owned by any one contributing agency alone.
  • Are Longitudinal (capture data from the same population over multiple years); Individual Level (include data that is specific to individual people but may contain identifiable information or be anonymous); and Statewide (bring together and connect data or records from multiple state agencies).
  • Are supported in large part by a number of federal grants awarded to states to develop or improve their data systems in order to effectively measure the success of educational programs. Longitudinal data is needed for effective measurement.
  • Are designed to help school districts, schools, and teachers make informed, data-driven decisions to improve student learning.
  • Leverage stakeholders and partners of education, training, and employment programs to create a system which provides data to support the research and evaluation of programs to improve the outcomes of individuals provided service.

Data developments are fraught with challenges. Most state P20W and SLDS are custom built. They require technology and personnel investments to develop their systems from scratch—this can take years to materialize. Maintaining data systems over time creates financial strain and risks long-term durability of such a system. Examples of problems endemic to these developments:

  1. They are unable to realize the promise of connecting full ecosystem of data sets across agencies.
  2. They often have performance constraints due to an inability to keep up with the latest innovations in technology and data management.
  3. They may have significant security vulnerabilities.
  4. Data retooling of systems is expensive, and time-consuming. States rely on continuing infusion of federal funding, raising issues around funding reliability over time.

Typically describes an individual’s demographics such as age, gender, race/ethnicity, current employment status, salary and wages. Personal data can also include an individual’s professional and academic goals.

As defined by the Velocity Network Foundation, in decentralized identity systems, a presentation is the act of showing credentials to a verifier. It involves selectively disclosing certain parts of a credential to prove authenticity without revealing more information than necessary, ensuring privacy-preserving verifications. 

According to the UN Refugee Agency, a refugee is a person forced to flee their country because of war, violence, or persecution. Persecution is founded in fear of persecution for reasons of race, religion, nationality, political opinion, or membership in a particular social group. War and ethnic, tribal and religious violence are leading causes of refugees fleeing their countries.

The 1951 Geneva Convention is the main international instrument of refugee law. The Convention defines who a refugee is and the types of legal protection and other assistance and social rights they should receive from the countries who have signed the document. The Convention was limited to protecting mainly European refugees in the aftermath of World War II, but a newer document, the 1967 Protocol, expanded the scope as the problem of displacement spread around the world.

Related terms: Internally displaced person, stateless person, asylum seeker

Refer to a series of AI models that teach robots to complete basic tasks in environments they have never been trained for, without additional training or fine-tuning. RUMs allow machines to complete five different tasks: opening doors and drawers, and picking up tissues, bags, and cylindrical objects in unfamiliar environments. This is a significant advance since researchers typically need to train robots on new data for each new environment they encounter — often a time-consuming, expensive process.

Structured data is organized and easily searchable, making it suitable for analysis and decision-making. Structured data typically has a pre-defined data model or fixed schema, such as structured rows and columns that can be sorted. Examples: Excel files; SQL databases; Web form results; Search Engine Optimization (SEO) tags; product directories; reservation systems.

Unstructured Data does not have a pre-defined data model or fixed schema, making it more difficult to organize and process using traditional data management tools. Examples: PDF; Word files; printed/scanned documents; Image, video, audio.

Refers to information that is artificially created rather than collected from real people or events. It is generated by computer programs to mimic the patterns and characteristics of real data without exposing anyone’s personal details. For example, instead of using actual student records to test a new education tool, developers might use synthetic data that looks and behaves like real student records but does not belong to any actual person. This allows organizations to test, train, and improve systems while protecting privacy and reducing risks.

The term simulated data is sometimes used in connection with synthetic data, especially when the information is produced through a computer simulation. In practice, all simulated data is synthetic, but not all synthetic data comes from simulations—some is generated by machine learning models, statistical methods, or rule-based systems.

As defined by the Velocity Network Foundation, represents the next evolution of the internet, characterized by decentralized technologies like blockchain, peer-to-peer networks, and distributed data storage. In contrast to Web 2.0, where centralized platforms like social media, cloud services, and big tech companies dominate the internet, Web 3.0 seeks to put more control in the hands of individual users, promoting ownership, privacy, and security through decentralization.

Web scraping is a technology tool to collect information from websites automatically. Instead of copying and pasting data by hand, a computer program quickly gathers the information and organizes it. The practice of web scraping is growing because the internet has more data than ever and collecting information in this way is used for many things:

  • Tracking Prices – websites can check prices on other websites to offer better deals.
  • Research – scientists and businesses collect data to study trends.
  • News & Alerts – companies and journalists monitor news websites for important updates and useful information.
  • Governments – monitor data for public services.
  • Job Hunting – some sites collect job postings from many sources.

Some legal restrictions can impact web scraping:

  • Some websites do not allow web scraping in their rules.
  • Copyright Laws – copying and using data without permission can be illegal.
  • Privacy Laws – scraping personal information (like emails or phone numbers) without consent is often against the law.

While web scraping is a growing practice, there is a growing role too for information aggregators that gather content from multiple sources to make it easier for users to find and understand. Web scraping plays a key role in this process by:

  • Automating Data Collection – instead of manually searching for new reports, projects, or glossary terms, web scraping can pull updates from various sources.
  • Keeping Information Current – aggregator databases need the latest data, and scraping can help refresh content quickly.
  • Improving Search & Navigation – scraping can help structure data so users can search, filter, and explore information more easily.

A difference between web scraping and information aggregators is that aggregator databases often combine automated methods (e.g., API integrations) with human review to maintain data accuracy, trustworthiness, and ethical standards.

SeeInformation Aggregator Database | Learn & Work Ecosystem Library

 

Refers to AI-generated audio, video, images, or digital identities used to impersonate employees, job applicants, executives, customers, or other individuals within employment and workforce settings. These synthetic media can be used for legitimate purposes such as training simulations or accessibility, but they are more commonly discussed in relation to fraud, cybersecurity threats, identity theft, hiring deception, financial scams, and workplace misinformation.

Examples include:

  • fake job candidates using AI-generated identities during virtual interviews
  • criminals impersonating executives to authorize financial transactions
  • fabricated employee communications
  • manipulated videos designed to damage reputations or spread false information.

As generative AI becomes more sophisticated, organizations are adopting stronger identity verification, authentication procedures, cybersecurity training, and governance policies to detect and reduce the risks associated with workforce deepfakes.

Organizations (521)

Initiatives (668)

Topic Briefs (159)