sbd.org.uk
Back to blog
Abstract geometric visualisation of Microsoft 365 tenant health monitoring
Paul

Paul

Solution Architect

···11 min read

M365 Tenant Health: Why It's Probably a Mess

Microsoft 365 tenant health audit checklist: detect orphaned groups, expired app secrets, CA policy sprawl, and SharePoint chaos with Graph API scripts.

microsoft-365entra-idgovernancepowershellgraph-apiautomationtenant-healthconditional-access

Every Microsoft 365 tenant I inherit has the same problems. The symptoms vary in severity, but the pattern is remarkably consistent: groups that nobody owns, SharePoint sites that nobody maintains, app registrations with secrets that expired six months ago, and a Conditional Access policy set that reads like geological strata from successive security audits.

These are not skills failures. The people who built and maintained these tenants were perfectly competent. The problem is that M365 tenants drift towards chaos by default. Every time someone creates a Teams channel, spins up a SharePoint site, or registers an application, the tenant gets a little messier, and without deliberate, automated governance, nothing reverses that drift.

This post is a tenant health assessment checklist. For each anti-pattern, I will show you how to detect it with Graph API, give you a sense of scale so you know whether your numbers are normal, and explain what to do about it. If you are looking for the broader case for standardisation and repeatable baselines, the Standardisation post covers that ground. This post is the diagnostic that tells you where to start.

Prerequisites: all scripts use the Microsoft Graph PowerShell SDK v2.x and require PowerShell 7+. Connect with the scopes you need before running each section. The snippets below are kept concise for readability; for production-ready Graph API patterns with proper error handling and parameterisation, see Microsoft Graph API for Architects.

Quick reference: tenant health metrics

IssueTypical Scale (2,000-user tenant)Detection Method
Ownerless Groups150-300 groupsGet-MgGroup + Get-MgGroupOwner
Dormant SharePoint Sites20-40% of all sitesSharePoint usage reports via Graph
Expired App Credentials30-50% of app registrationsGet-MgApplication + PasswordCredentials
Non-Compliant Naming80%+ without enforced policyRegex pattern match
Redundant CA Policies10-30% overlapGet-MgIdentityConditionalAccessPolicy

The scripts across every section need these read-only permissions: Group.Read.All, Application.Read.All, Policy.Read.All, Reports.Read.All.


1. Orphaned M365 groups

M365 groups accumulate faster than anyone expects. Every Teams team creates one. Every Planner board creates one. So does every Yammer community. People create groups freely, then leave the organisation without transferring ownership. The group persists, ownerless, with members who have no idea they are in it and content that may or may not still be relevant.

In a 2,000-user tenant, expect 150 to 300 ownerless groups. Tenants that have been running for five years or more without governance are worse. I have seen 10,000-user tenants with 3,000+ groups where fewer than half had an identifiable owner.

Detecting ownerless groups

Connect-MgGraph -Scopes "Group.Read.All"
 
# Get all Microsoft 365 groups
$groups = Get-MgGroup -Filter "groupTypes/any(g:g eq 'Unified')" -All `
    -Property Id,DisplayName,CreatedDateTime,Description
 
$ownerlessGroups = foreach ($group in $groups) {
    $owners = Get-MgGroupOwner -GroupId $group.Id
    if ($owners.Count -eq 0) {
        [PSCustomObject]@{
            DisplayName = $group.DisplayName
            Created     = $group.CreatedDateTime
            Description = $group.Description
        }
    }
}
 
Write-Output "Ownerless groups: $($ownerlessGroups.Count) of $($groups.Count) total"
$ownerlessGroups | Sort-Object Created | Format-Table -AutoSize

Finding stale groups

Ownerless is bad. Ownerless and inactive is a cleanup candidate.

# M365 Groups activity report (requires Reports.Read.All)
$uri = "https://graph.microsoft.com/v1.0/reports/getOffice365GroupsActivityDetail(period='D90')"
Invoke-MgGraphRequest -Uri $uri -OutputFilePath "$env:TEMP\GroupActivity.csv"
$report = Import-Csv "$env:TEMP\GroupActivity.csv"
 
$staleGroups = $report | Where-Object {
    $_.'Last Activity Date' -eq '' -or
    ([datetime]$_.'Last Activity Date') -lt (Get-Date).AddDays(-90)
}
 
Write-Output "Groups with no activity in 90 days: $($staleGroups.Count)"

What to do about it

In the short term, assign owners to critical groups manually. For the rest, contact the last known owner or the most active member and ask them to take ownership or confirm the group can be deleted.

In the long term, enable the M365 group expiration policy in Entra ID (Groups > Expiration) and set a lifecycle of 180 or 365 days. Groups with activity renew automatically, and groups without owners receive renewal reminders sent to configured fallback admins. A group that is not renewed is deleted, and its owners or an administrator can restore it within 30 days of deletion. This requires Entra ID P1 or higher.

You should also restrict who can create M365 groups. By default, every licensed user can create them, which is how you end up with hundreds of orphans. Limit creation to a security group of people who understand the naming conventions and governance expectations. Configure this via the GroupCreationAllowedGroupId directory setting.


2. SharePoint site sprawl and dormant sites

Every M365 group gets a SharePoint site. Every Teams team gets one. People also create standalone communication sites and classic sites. Nobody tracks them centrally. You end up with "Marketing", "Marketing Team", "Marketing-old", "Marketing 2024", and "Marketing (DO NOT USE)".

SharePoint sites pile up too. A 2,000-user tenant commonly runs 500 to 1,500 of them. Of those, 20-40% typically have no identifiable owner or the listed owner has left the organisation.

Detecting dormant sites

Connect-MgGraph -Scopes "Reports.Read.All"
 
$uri = "https://graph.microsoft.com/v1.0/reports/getSharePointSiteUsageDetail(period='D90')"
Invoke-MgGraphRequest -Uri $uri -OutputFilePath "$env:TEMP\SharePointUsage.csv"
$sites = Import-Csv "$env:TEMP\SharePointUsage.csv"
 
$dormantSites = $sites | Where-Object {
    $_.'Last Activity Date' -eq '' -or
    ([datetime]$_.'Last Activity Date') -lt (Get-Date).AddDays(-90)
}
 
Write-Output "Dormant sites (no activity in 90 days): $($dormantSites.Count) of $($sites.Count)"
 
# Identify potential duplicates by normalised name
$siteNames = $sites | ForEach-Object { ($_.'Site URL' -split '/')[-1] }
$potentialDuplicates = $siteNames |
    Group-Object { $_ -replace '[-_\d]','' } |
    Where-Object { $_.Count -gt 1 }
 
Write-Output "`nPotential duplicate site groups:"
$potentialDuplicates | ForEach-Object {
    Write-Output "  $($_.Name): $($_.Group -join ', ')"
}

The duplicate-detection step above depends on real site names, and Microsoft 365 usage reports conceal them by default. Clear "Conceal user, group, and site names in all reports" under Settings, Org Settings, Services tab, Reports before you rely on this script.

What to do about it

Reassign ownership for sites where the listed owner has left. This is non-negotiable, because an unowned site has nobody accountable for it.

Implement an inactive site policy via the SharePoint admin centre. SharePoint now supports native inactive site policies that automatically notify owners when their site has been dormant and can restrict access after a defined period. This is far more practical than manually reviewing hundreds of sites quarterly.

Apply a naming convention too. More on this in section 4, but the fix for "Marketing (DO NOT USE)" is a naming standard enforced at creation time, not hoped for after the fact.


3. Entra app registrations with expired secrets

This one is insidious because it fails silently. A developer creates an app registration, adds a client secret with a two-year expiry, and moves on. The integration works perfectly for 23 months. Then the secret expires on a Saturday night and an automated process stops working. Nobody knows who created the app, what it does, or where the secret is used.

App registrations are fewer in absolute terms, 50 to 200 in a 2,000-user tenant. In my experience, 30-50% have expired credentials and 20-40% have no assigned owner.

It is also attack surface rather than only technical debt. An unowned app registration with a stale secret is exactly the kind of unwatched object that directory enumeration surfaces first, and any member of your tenant can list every one of them by default.

The credential audit

Connect-MgGraph -Scopes "Application.Read.All"
 
$apps = Get-MgApplication -All `
    -Property Id,AppId,DisplayName,PasswordCredentials,KeyCredentials
$now = Get-Date
$warningThreshold = $now.AddDays(30)
 
$credentialReport = foreach ($app in $apps) {
    $owners = Get-MgApplicationOwner -ApplicationId $app.Id
    $ownerNames = if ($owners) {
        ($owners.AdditionalProperties.displayName -join '; ')
    } else { "NO OWNER" }
 
    foreach ($secret in $app.PasswordCredentials) {
        [PSCustomObject]@{
            AppName   = $app.DisplayName
            AppId     = $app.AppId
            Type      = "Secret"
            Expires   = $secret.EndDateTime
            Status    = if ($secret.EndDateTime -lt $now) { "EXPIRED" }
                       elseif ($secret.EndDateTime -lt $warningThreshold) { "EXPIRING SOON" }
                       else { "OK" }
            Owners    = $ownerNames
        }
    }
 
    foreach ($cert in $app.KeyCredentials) {
        [PSCustomObject]@{
            AppName   = $app.DisplayName
            AppId     = $app.AppId
            Type      = "Certificate"
            Expires   = $cert.EndDateTime
            Status    = if ($cert.EndDateTime -lt $now) { "EXPIRED" }
                       elseif ($cert.EndDateTime -lt $warningThreshold) { "EXPIRING SOON" }
                       else { "OK" }
            Owners    = $ownerNames
        }
    }
}
 
$expired = ($credentialReport | Where-Object Status -eq "EXPIRED").Count
$expiringSoon = ($credentialReport | Where-Object Status -eq "EXPIRING SOON").Count
$noOwner = ($credentialReport | Where-Object Owners -eq "NO OWNER" |
    Select-Object AppName -Unique).Count
 
Write-Output "Expired credentials: $expired"
Write-Output "Expiring within 30 days: $expiringSoon"
Write-Output "Apps with no owner: $noOwner"

What to do about it

Assign an owner to every app registration. Entra ID sends expiration notifications to app owners at 60, 30, 15, and 7 days before credential expiry, but only if an owner is listed. No owner means no warning.

Prefer certificates over client secrets. Certificates are harder to leak, because you cannot copy them from the portal, and they support longer validity and align better with zero-trust principles.

Prefer managed identities where possible. If the workload runs in Azure, in Logic Apps, Azure Functions or Automation Accounts, use a managed identity instead of an app registration. There is no credential to manage, monitor or rotate. This is the most effective way to reduce app registration sprawl.

Enforce tenant-wide credential policies too. The application authentication methods policy lets you restrict secret lifetime across the tenant, block password credentials entirely, and force certificates. If you cannot justify a blanket restriction, at minimum set a maximum secret lifetime of 180 days, so the longest possible gap between a secret being created and expiring with nobody watching is six months rather than two years.


4. Naming convention chaos

Without an enforced naming policy, you get Teams called "Project Alpha", groups called "prj-alpha-team", and SharePoint sites called "Alpha Project Site". Search becomes guesswork. Departments create duplicate resources because they cannot find existing ones.

This is almost universal. In tenants without enforced naming policies, expect inconsistent naming in 80%+ of groups, sites, and Teams.

Measuring compliance

Connect-MgGraph -Scopes "Group.Read.All"
 
# Define your expected naming pattern
# Example: department prefix like "IT-", "HR-", "FIN-"
$namingPattern = '^(IT|HR|FIN|MKT|ENG|OPS|SEC|LEGAL)-'
 
$groups = Get-MgGroup -Filter "groupTypes/any(g:g eq 'Unified')" -All `
    -Property DisplayName,CreatedDateTime
 
$compliant = ($groups | Where-Object { $_.DisplayName -match $namingPattern }).Count
$total = $groups.Count
 
Write-Output "Naming compliance: $compliant/$total ($([math]::Round($compliant/$total*100,1))%)"

What to do about it

Enable the Entra ID group naming policy. It requires P1, and it enforces prefixes and suffixes automatically at creation time. You can use attribute-based prefixes, for example [Department]-[GroupName], and block offensive or reserved words. Configure it in the Entra admin centre under Groups > Naming policy.

The important caveat is that the naming policy only applies to groups created after the policy is enabled. Existing groups need to be renamed manually or via script. This is a one-time effort, and worth doing. A tenant where every group follows a consistent naming convention is far easier to administer than one where you have to guess what "Project Phoenix" refers to.


5. Conditional Access policy sprawl

This is the most dangerous anti-pattern on the list. Organisations accumulate Conditional Access policies over years. Each security incident or compliance audit produces a new policy. Nobody removes old ones. Policies overlap, contradict each other, and interact in unexpected ways. The "What If" tool in Entra ID exists specifically because nobody can reason about their CA policy set by reading the policies alone.

CA policies are the smallest number on this list: 15 to 40 in a 2,000-user tenant. Tenants that have been through multiple security audits can have 60+. Of those, 10-30% are often redundant or superseded by newer policies.

For a detailed guide to building a CA policy set from scratch (including the persona-based framework, break-glass accounts, and the evaluation flow), see the Conditional Access post. This section focuses on auditing what you already have.

The CA audit

Connect-MgGraph -Scopes "Policy.Read.All"
 
$policies = Get-MgIdentityConditionalAccessPolicy -All
 
$caReport = foreach ($policy in $policies) {
    [PSCustomObject]@{
        Name          = $policy.DisplayName
        State         = $policy.State
        Modified      = $policy.ModifiedDateTime
        IncludeUsers  = ($policy.Conditions.Users.IncludeUsers -join '; ')
        IncludeApps   = ($policy.Conditions.Applications.IncludeApplications -join '; ')
        GrantControls = ($policy.GrantControls.BuiltInControls -join '; ')
    }
}
 
$enabled = ($caReport | Where-Object State -eq 'enabled').Count
$reportOnly = ($caReport | Where-Object State -eq 'enabledForReportingButNotEnforced').Count
$disabled = ($caReport | Where-Object State -eq 'disabled').Count
 
Write-Output "CA policies: $($policies.Count) total"
Write-Output "  Enabled: $enabled"
Write-Output "  Report-only: $reportOnly"
Write-Output "  Disabled: $disabled (review these for removal)"
 
# Flag policies not modified in over a year
$staleDate = (Get-Date).AddYears(-1)
$stalePolicies = $caReport | Where-Object {
    $_.Modified -and ([datetime]$_.Modified) -lt $staleDate
}
Write-Output "`nPolicies not modified in over 1 year: $($stalePolicies.Count)"
$stalePolicies | Select-Object Name, State, Modified | Format-Table -AutoSize

Auditing named locations

Named locations are the forgotten dependency. Policies reference them, but nobody maintains the IP ranges. People leave, offices close, VPN endpoints change, and the named location still says "London Office" with a CIDR block that now belongs to a different tenant.

# Find named locations not used by any policy
$locations = Get-MgIdentityConditionalAccessNamedLocation -All
 
$unusedLocations = foreach ($loc in $locations) {
    $usedBy = $policies | Where-Object {
        $_.Conditions.Locations.IncludeLocations -contains $loc.Id -or
        $_.Conditions.Locations.ExcludeLocations -contains $loc.Id
    }
    if ($usedBy.Count -eq 0) {
        [PSCustomObject]@{
            Name = $loc.DisplayName
            Type = $loc.AdditionalProperties.'@odata.type' -replace '#microsoft.graph.',''
        }
    }
}
 
Write-Output "Named locations not referenced by any policy: $($unusedLocations.Count)"
$unusedLocations | Format-Table -AutoSize

What to do about it

Export and map everything. Get all policies into a spreadsheet, document what each one does and why it exists, and identify overlaps. Use the What If tool to test common scenarios, such as a standard user signing in from a managed device, from a personal device, and from overseas.

Adopt a persona-based framework. Define four to six user personas, such as standard users, privileged admins, guests, service accounts and break-glass accounts, and give each a clear, documented policy set. If you cannot explain which policies apply to a given persona without running What If, your policy set is too complex.

Remove disabled policies. A disabled CA policy is not a safety net: it is clutter that makes the active policy set harder to understand. If you disabled it for a reason, document that reason and delete the policy. If you might need it again, export the JSON first.

Verify break-glass exclusions. Exclude at least two break-glass accounts from every CA policy, and make them cloud-only, protected with FIDO2 keys stored physically in a safe, and monitored with alerts on any sign-in. If even one CA policy does not exclude your break-glass accounts, you have a lockout risk.


6. Security defaults vs. custom CA confusion

More tenants get this wrong than you would expect. Security defaults provide baseline MFA and block legacy authentication. They are designed for organisations without Entra ID P1 or P2. Custom Conditional Access policies require P1. The two are mutually exclusive: you cannot turn both on at the same time.

Microsoft blocks having both enabled at once, so that is not the risk; the real risk is disabling security defaults to implement CA policies and then failing to replicate the protections security defaults provided.

Checking for gaps

Connect-MgGraph -Scopes "Policy.Read.All"
 
$secDefaults = Invoke-MgGraphRequest `
    -Uri "https://graph.microsoft.com/v1.0/policies/identitySecurityDefaultsEnforcementPolicy"
 
Write-Output "Security Defaults enabled: $($secDefaults.isEnabled)"
 
if (-not $secDefaults.isEnabled) {
    $policies = Get-MgIdentityConditionalAccessPolicy -All |
        Where-Object { $_.State -eq 'enabled' }
 
    $mfaForAll = $policies | Where-Object {
        $_.Conditions.Users.IncludeUsers -contains 'All' -and
        $_.GrantControls.BuiltInControls -contains 'mfa'
    }
 
    $blockLegacy = $policies | Where-Object {
        ($_.Conditions.ClientAppTypes -contains 'exchangeActiveSync' -or
         $_.Conditions.ClientAppTypes -contains 'other') -and
        $_.GrantControls.BuiltInControls -contains 'block'
    }
 
    $gaps = @()
    if (-not $mfaForAll) { $gaps += "No policy requiring MFA for all users" }
    if (-not $blockLegacy) { $gaps += "No policy blocking legacy authentication" }
 
    if ($gaps.Count -gt 0) {
        Write-Warning "Security Defaults are OFF but CA policies have gaps:"
        $gaps | ForEach-Object { Write-Warning "  - $_" }
    } else {
        Write-Output "CA policies appear to cover Security Defaults protections."
    }
}

What to do about it

If you have Entra ID P1 or P2, disable security defaults and implement CA policies that cover the protections security defaults provided: MFA registration for all users, risk-based MFA enforcement, require MFA for admin roles, block legacy authentication, and protect Azure management endpoints.

If you do not have P1 or P2, keep security defaults enabled and do not try to use CA policies.

Either way, document the decision. This is a common audit finding, and assessors want to see that the choice was deliberate rather than accidental.


The governance stack that prevents all of this

Detecting these problems is the easy part; actual governance means preventing them from recurring.

Microsoft has shipped most of the tooling you need in the last two years, but none of it is enabled by default.

Group lifecycle

CapabilityWhat it does
Expiration policyAuto-deletes groups with no owner and no activity after a configurable period of 180 or 365 days. Requires Entra ID P1.
Creation restrictionsLimits M365 group creation to a security group, and stops the unchecked proliferation at the source.
Naming policyEnforces department prefixes and blocked words. Requires P1.
Access reviewsPeriodically prompts group owners to confirm memberships are still correct, with multi-stage escalation if the owner does not respond. The base access review capabilities need Entra ID P2, and the newer ones need Entra ID Governance or Entra Suite.

Application governance

CapabilityWhat it does
Authentication methods policyRestricts secret lifetime, blocks password credentials, and forces certificates.
Owner assignment enforcementMakes it a policy that no app registration ships without at least two owners. Entra sends expiry notifications to owners automatically.
Workload identity recommendationsPart of Workload Identities Premium, a separate licence from Entra ID P1 or P2. Surfaces unused apps, overprivileged service principals, and credential expiry warnings.
Managed identitiesUse them for every Azure-hosted workload: no secrets, no expiry, no rotation.

Conditional Access hygiene

CapabilityWhat it does
Authentication strengthsReplace the blunt require MFA grant control with fine-grained requirements, such as phishing-resistant only or passwordless only. They are now generally available, and because an authentication strength is a Conditional Access grant control, every strength needs Entra ID P1, not only custom ones.
Policy templatesPre-built policies for common scenarios, such as blocking legacy authentication or requiring MFA for admins. Use them as a starting point rather than building from scratch.
Regular review cadenceReview your CA policy set quarterly. If a policy has been in report-only mode for more than three months, either promote it to enforced or delete it.

Running the full audit

Each section above stands alone, but the real value is running all the checks together and producing a single tenant health report. The workflow is straightforward:

  1. Create an Entra ID app registration with the read-only permissions from each section (Group.Read.All, Application.Read.All, Policy.Read.All, Reports.Read.All).
  2. Run each detection script in sequence.
  3. Export everything to a single workbook with one sheet per anti-pattern.
  4. Present the findings with estimated impact: number of ownerless groups, dormant sites, expired credentials, non-compliant names, and CA policy gaps.

In my experience, the first audit always surfaces enough issues to justify the time it took to run, and usually several times over. The Licensing Audit post covers the financial side, orphaned licences and E5-to-E3 downgrade candidates. This post covers the structural side. Together, they cover both sides of tenant health.

Implementing the governance stack and keeping it maintained is harder than running the audit, and it takes executive sponsorship, documented standards, and automation that runs whether anyone remembers to check or not.


If you have inherited a particularly messy tenant and found creative ways to clean it up, I would be interested to hear what worked. The scripts are the easy part; the organisational change management is where the real war stories live.

Get the next one by email

Long, specific write-ups on M365 architecture and security, worked out against real tenants rather than summarised from documentation. Sent rarely, and only when it is worth your time.