Netwrix 1Secure delivers unified visibility across data and identity - free for 14 days with full access. Start a free trial

Resource centerBlog

How to find and secure unknown data assets

How to find and secure unknown data assets

Oct 6, 2026

Unknown data assets create inventory gaps that weaken security, compliance, and cyber resilience, since untracked stores sit outside classification, access governance, and continuous visibility. Teams can't assess sensitive content, verify effective access, or keep their data security posture current as assets and permissions change. Continuous discovery, classification, and access reviews close the gap.

Most data security programs can protect only the assets they know exist. Shares, buckets, and exports outside the inventory also fall outside classification, access governance, and continuous monitoring, creating blind spots where sensitive data can remain exposed longer.

Those blind spots show up in breach data. Shadow data appeared in 35% of breaches analyzed in IBM's Cost of a Data Breach Report 2024, the most recent edition that breaks out shadow data. Breaches involving shadow data also took longer to identify and contain, with an average lifecycle of 291 days and an average cost of $5.27 million. An incomplete inventory undermines detection and response.

That’s why the National Institute of Standards and Technology Cybersecurity Framework (NIST CSF) 2.0 emphasizes maintaining data inventories and corresponding metadata.

Those inventories provide the foundation for controls such as DLP, classification, access reviews, and continuous monitoring. When a data store is missing from the inventory, those controls are far less likely to cover it.

What are unknown data assets?

An unknown data asset is any data store absent from an organization's current inventory, such as a share that outlived its project, a bucket in a secondary cloud account, or an export a departed analyst left behind. The term covers dark data and shadow data, but neither label matters, since an asset is worth investigating if nobody’s currently accounting for it.

Term

What creates it

Where it typically lives

Urgency signal

Dark data

Data collected but never analyzed or acted on

Cloud buckets, legacy shares, SaaS exports, AI caches, log repositories

Lower until classification shows sensitive content

Shadow data

Someone copies, backs up, replicates, or exports data outside centralized governance

Personal cloud accounts, unsanctioned SaaS, unmanaged app connections

Higher; the data sits outside centralized governance and monitoring

Unknown data assets

This guide's operational term for any asset absent from the current inventory, for any reason

Anywhere: on-premises shares, cloud storage, backups, forgotten SaaS repositories

Depends on content; the inventory gap defines the category

Why unknown data assets keep piling up

New assets arrive faster than any manual inventory records them, so the gap persists after a point-in-time review rather than closing.

  • Employee-provisioned storage: Staff open their own cloud storage accounts when the sanctioned option is slower or more restrictive. Netskope Threat Labs' Cloud and Threat Report: 2026 found that 31% of users in the average organization upload data to personal cloud apps every month, and that 60% of insider threat incidents involve personal cloud app instances.
  • SaaS sprawl: SaaS adoption compounds broader technology sprawl. Forrester's Q2 2024 Tech Pulse Survey found that 77% of US technology decision-makers report moderate to extensive technology sprawl.
  • Orphaned systems: Application decommissioning can leave backups and exports that outlive the source system, with no owner left to account for them. Cloud platforms make this the default behavior, since manual Amazon Relational Database Service (RDS) snapshots survive deletion of the source instance.
  • Shadow AI adoption: Generative AI (GenAI) tools copy sensitive data when someone pastes it into a prompt or connects a tool to a data source. A Gartner survey of 175 employees, published February 5, 2026, found that over 57% had used personal GenAI accounts for work and 33% admitted entering sensitive information into unapproved tools.

Why the inventory gap matters more than most security teams assume

Inventory scope determines which assets each data security control can reach. Exposure beyond that scope remains unmeasured.

No control reaches an asset nobody knows exists

DLP, SIEM, access governance, and encryption only cover the stores someone points them at. A repository nobody catalogs or classifies falls outside control coverage, reducing visibility, widening unmanaged exposure, and slowing incident-response readiness.

Most organizations know about this gap and haven't closed it

The Netwrix 2026 Data and Identity Security Report surveyed 2,317 IT and security leaders and produced stark findings. 55% of organizations don't continuously maintain a sensitive data inventory, and 74% can't get a single, unified view of where their sensitive data resides and which identities can access it. Continuous inventory maintenance and unified access visibility remain uncommon.

Unknown assets disproportionately show up in breach forensics

The Unit 42 2025 Global Incident Response Report names unmanaged and unmonitored assets, including endpoints, applications, and shadow IT, as easy entry points for attackers.

McLeod Health found a suspicious file on a server mid-decommissioning on March 5, 2026, according to its incident notice, 137 days after the unauthorized access on October 17-18, 2025. Continuous asset visibility supports faster investigation and stronger cyber resilience when those systems surface during incident response.

Netwrix DSPM finds sensitive data across on-premises, cloud, and SaaS environments and prioritizes access risk so security teams act on what matters most. Book a demo.

A repeatable method for finding unknown data assets

Each environment has its own enumeration tools and blind spots, so discovery should run per environment to account for those differences.

Get read-only access to every environment first

Discovery only finds what its account can see, so request read-only access to each environment before you run a single enumeration. Use named, dedicated identities rather than shared admin credentials, so every discovery query is attributable.

  • On-premises: Read access to Active Directory, a remote-management path to each server, and a login on each SQL Server instance you query. The workflow below reaches remote servers through CIM sessions.
  • AWS: Sign in to the management account or a delegated administrator to create an organization-wide AWS Config aggregator, and attach the AWSConfigRoleForOrganizations managed policy to its role.
  • Azure: Grant read access at the management group level. Azure Resource Graph returns no results for resources the account can't read, so a missing grant can make part of the estate look empty.
  • Google Cloud: Grant Cloud Asset Viewer (roles/cloudasset.viewer) on the organization, folder, or project you search.
  • Entra ID and SaaS: Listing delegated OAuth grants through Microsoft Graph takes the Directory.Read.All permission, and a signed-in admin needs a role such as Global Reader.
  • Backups and snapshots: These sit behind the backup console or the cloud snapshot APIs, not the source system, so include them in the access request.
  • Unregistered apps: Apps nobody registered have no admin record to read. Feed Cloud Discovery from Defender for Endpoint, a log collector, or uploaded firewall and proxy logs instead.

Anything a read-only account can't reach is itself a finding. Log it as an access gap and escalate it for a decision.

On-premises servers, shares, and databases

Start from Active Directory and work outward, since it's the fastest path to a complete list of on-premises systems to check.

  1. Pull file-server hostnames from Active Directory with Get-ADComputer, following the Get-ADComputer reference.
  2. Enumerate Server Message Block (SMB) shares over Common Information Model (CIM) sessions with Get-SmbShare -Special $false to exclude administrative shares.
  3. Pull sys.servers in SQL Server to return every linked or remote server, which often points to instances nobody documented.

Run steps 1 and 2 in one pass with this PowerShell workflow:

      $servers = Get-ADComputer -Filter * -Properties OperatingSystem |
    Where-Object OperatingSystem -Like '*Server*'
foreach ($server in $servers) {
    $session = New-CimSession -ComputerName $server.DNSHostName
    Get-SmbShare -CimSession $session -Special $false
    Remove-CimSession $session
}
      

The output’s a full share and server inventory to compare against whatever's already documented; anything missing from that documentation is an unknown asset.

Cloud storage, backups, and databases

Each provider keeps its own inventory tools, so enumeration must run separately in each one before you can compare the results.

  1. Enumerate storage resources at the organization level in each provider in use. AWS Config configuration aggregators cover Amazon Web Services; an Azure Resource Graph query for Microsoft.Storage/storageAccounts at management group scope covers up to 10,000 Microsoft Azure subscriptions, and Cloud Asset Inventory searches the storage.googleapis.com/Bucket asset type for Google Cloud.
  2. Run the Elastic Block Store (EBS) AWSSupport-AnalyzeEBSResourceUsage runbook in AWS to list volumes in available state and snapshots whose source volume no longer exists.
  3. Check Business Continuity Center in Azure for deprovisioned recovery points left behind after users deprovision their source resources.
  4. Compare the combined output against the current inventory, and treat any storage account, volume, or recovery point without a named, active project as a review candidate, since native tools only see resources someone has enrolled.

Any result absent from the current inventory is an unknown asset by definition, whether or not the source system still runs. In 2025, security researcher Jeremiah Fowler found 378 gigabytes of Navy Federal Credit Union backup files exposed exactly this way, sitting in a public Amazon S3 bucket nobody was tracking.

SaaS applications, OAuth grants, and shadow AI

Entra ID and governance dashboards only show apps someone has already registered or consented to, so finding the ones nobody registered takes a separate, traffic-based step.

  1. Run Defender for Cloud Apps Cloud Discovery against firewall, proxy, or endpoint traffic logs to surface every SaaS app in use, including ones with no Open Authorization (OAuth) grant or admin record.
  2. Filter those Cloud Discovery results to Microsoft's generative AI app category to isolate unsanctioned AI tools specifically.
  3. Feed the Purview browser extension for Edge and Chrome into the Insider Risk Management Risky AI usage template, which detects prompts and responses containing sensitive information on onboarded devices and catches content Cloud Discovery's traffic logs can't see.
  4. Review external sharing links separately, since the standard SharePoint sharing report excludes Anyone links; the SharePoint admin center's Data Access Governance reports include them, but require the applicable SharePoint management add-on and cover a 28-day window.
  5. Before any sanctioned AI rollout, run the required SharePoint management add-on's Everyone Except External Users (EEEU) report, which lists the top 100 sites shared with the entire organization in the past 28 days. Every one of those sites becomes AI-searchable the moment Copilot goes live, since Copilot grounds responses in data users already have permission to access.

How to tell which unknown assets matter

Not every untracked asset deserves the same urgency; content and access should set the priority order, regardless of how obscure the asset's location is.

Prioritize content over location

An asset containing personally identifiable information (PII), protected health information (PHI), or payment card data is urgent, no matter how obscure its path; a duplicate of non-sensitive data can wait.

NIST SP 800-122 lists six confidentiality impact factors and warns that they interact, since one factor alone might indicate a low impact level while another overrides it toward high impact. Review the complete NIST impact guidance for the full framework.

Classify before ruling anything out as noise

Use pattern-based and contextual classification to separate real exposure from noise. Purview confidence levels for sensitive information types run 65, 75, or 85; the low setting catches the most matches and the most false positives.

Known test primary account numbers such as Visa's 4111111111111111 pass Luhn validation, so exclude them through allow lists or exact-match reference tables before they consume remediation time.

Flag broad access plus sensitive content first

NIST SP 800-122 explains that more people and systems accessing PII create more opportunities to compromise its confidentiality. Treat sensitive data reachable by Everyone, Authenticated Users, or an open sharing link as the first remediation tier, ahead of sensitive data with tightly scoped access.

How to secure the unknown data assets you find

The Netwrix 2026 report found that 75% of incident-based data exposures begin with compromised identities or misconfigured permissions, so identity carries every step that follows.

Fix access before anything else

For every newly discovered asset with confirmed sensitive data or broad exposure, contain excessive access first; if the content's sensitivity is still unknown, classify it promptly to determine the final remediation priority.

Permissions nobody reviews keep the asset exposed, and the Verizon 2026 Data Breach Investigations Report found that half of permission misconfiguration findings took almost eight months to resolve.

Strip Everyone and Authenticated Users from ACLs, expire Anyone links (CISA's ScubaGear baseline sets the default sharing scope to "Specific people"), and apply least privilege under NIST AC-6 before ongoing monitoring starts.

Answer "who can reach this, and should they?" for every asset

Answering that question requires effective-access analysis, since raw ACLs don't show the complete result. A few identity paths need separate attention.

In Microsoft Entra ID, transitiveMemberOf flattens nested groups for users and service principals, but application assignments don't cascade to nested groups, so directory enumeration alone overstates who can open an app.

Any links exist entirely outside directory queries, so review sharing-link audits on their own, and include non-human identities in the review as well.

An asset isn’t secured until every human or non-human identity that can reach it has a name and a confirmed reason for access.

Put the asset under continuous monitoring

After inventorying the asset and correcting access, apply the same ongoing visibility used everywhere else, including file share audit events or the equivalent cloud audit trail, plus NIST CA-7 continuous monitoring and CM-3 change control so permissions don't drift back.

Shadow data escapes existing access controls and the tools that monitor and log data access, which makes "known but unmonitored" its own risk category, separate from "unknown." Continuous visibility turns both categories into measurable data security posture work.

How Netwrix helps find and secure unknown data assets

Most discovery efforts stall on the same three things: knowing what's actually sensitive, knowing who can reach it, and proving both when someone asks.

Closing the discovery-to-remediation gap

Netwrix DSPM provides the umbrella capability for finding and protecting sensitive data, prioritizing compliance risk, and addressing risky access across hybrid environments.

Netwrix Access Analyzer serves as its enterprise data security posture management engine, with data discovery, classification, Data Access Governance, and more than 40 data collection modules spanning file systems, SharePoint, databases, and cloud storage.

Turning raw permissions into effective access

Access Analyzer resolves nested group membership to show the access a user or account actually has, rather than the raw permissions listed on an object, so teams can review effective access directly instead of reconstructing it by hand.

Risk-based prioritization then directs remediation toward sensitive data that combines open or excessive access with confirmed sensitive content, following the same triage this guide recommends. Teams can assign data owners and track remediation decisions for the risks it surfaces, which turns a one-time cleanup into a repeatable governance process.

Proving the fix worked

An unexpected spike of 27,000 file changes hit a server holding regulated data at Cheshire County Government. Its five-person IT team traced the change to a permissions misconfiguration and closed the investigation in 15 minutes with Netwrix Auditor for Active Directory and Windows file servers, instead of the days a manual log review would have taken.

First National Bank Minnesota rebuilt its Active Directory environment to lock down income verification records, Social Security numbers, and employment history to a strict need-to-know basis. Netwrix Auditor showed exactly where that sensitive data lived and who could reach it, and the bank completed a rebuild it had budgeted at six months in three weeks.

Extending the same model across the estate

Access Analyzer supports discovery and classification across on-premises and cloud data sources, including Microsoft environments and Amazon S3, with Azure Files support as well, so the same classification and access-review model carries from file servers to cloud and SaaS repositories.

Cloud storage and SaaS OAuth connections are often the most defensible place to start, given how much of the app estate arrives as shadow IT; from there, the same review extends to on-premises shares, backups, and exports.

Close the inventory gap before it becomes a breach

An inventory gap doesn't stay a documentation problem for long. It becomes a security control gap the moment an unknown asset holds sensitive data or open access, and the cost by then shows up in breach timelines rather than spreadsheet rows.

Run discovery continuously across every environment, prioritize by content and access rather than convenience, and close the loop with monitoring so an asset that surfaces once doesn't go dark again.

Request a demo to see how Netwrix DSPM turns unknown data assets into a governed, continuously monitored part of the inventory.

Frequently asked questions about how to find and secure unknown data assets

Share on

Learn More

About the author

Asset Not Found

Netwrix Team