Technical series

Read by operational path, not only by date.

Naxaya articles are grouped into practical paths so a reader can move from context to diagnosis, implementation and runbook-level validation without hunting through the full archive.

Cloud 12 articles

Azure WAF operations

Read WAF blocks in KQL, qualify false positives, then add targeted OWASP/CRS exclusions or custom rules with evidence.

From noisy blocks to defensible WAF changes.

  1. 01 Azure WAF: read Application Gateway blocks with KQL without chasing every layer
  2. 02 WAF and KQL: identify a false positive before creating an exclusion
  3. 03 Azure WAF: add an OWASP/CRS exclusion without weakening all protection
  4. 04 Azure WAF: when to use custom rules before managed OWASP rules
  5. 05 Azure WAF: frame an emergency custom rule without losing evidence
  6. 06 Azure WAF: move a policy from Detection to Prevention without breaking traffic
  7. 07 Azure WAF: prepare an evidence pack before a policy PR
  8. 08 Azure WAF: diagnose rate limiting before increasing the threshold
  9. 09 Azure WAF: validate a managed rule upgrade before switching to Prevention
  10. 10 Azure WAF: diagnose a file upload before raising size limits
  11. 11 Azure WAF Bot Manager: qualify a bot before adding an allow rule
  12. 12 Azure WAF: diagnose inconsistent policy enforcement after deployment
Cloud 35 articles

Azure private networking

Separate private exposure, outbound networking, DNS, Load Balancer health, Private Endpoint validation and Application Gateway troubleshooting.

Make private Azure paths testable and explainable.

  1. 01 Azure Private Endpoint: build a validation matrix before production
  2. 02 Do not confuse private exposure, outbound networking and application security on Azure
  3. 03 Azure VNet Integration: diagnose outbound networking before changing the application
  4. 04 Azure App Service VNet Integration: diagnose subnet exhaustion before retrying scale
  5. 05 Azure NAT Gateway: diagnose SNAT exhaustion and outbound IP drift
  6. 06 Azure AKS: diagnose outbound SNAT exhaustion before adding a NAT Gateway
  7. 07 Azure Application Gateway: diagnose 502 errors without mixing DNS, TLS and backend health
  8. 08 Azure hybrid DNS: when to use Private Resolver, on-premises forwarders and private zones
  9. 09 Azure Private DNS Resolver: diagnose split-brain DNS before changing zones
  10. 10 Azure Private Endpoint: detect Terraform, DNS, and network drift before incident
  11. 11 Azure: make private paths verifiable with synthetic probes
  12. 12 Azure internal APIM: diagnose a private API before changing policies
  13. 13 Azure Container Apps: diagnose private ingress before changing revisions
  14. 14 Azure Container Apps: diagnose outbound egress before changing code
  15. 15 Azure AKS: diagnose private ingress before changing deployments
  16. 16 Azure Functions: diagnose a private HTTP endpoint before changing code
  17. 17 Azure App Service: diagnose a private endpoint before redeploying
  18. 18 Azure Storage: diagnose a private endpoint without opening the account
  19. 19 Azure SQL: diagnose a private endpoint before changing the database
  20. 20 Azure Service Bus: diagnose a private endpoint before touching queues
  21. 21 Azure UDR and NAT Gateway: diagnose egress before opening the firewall
  22. 22 Azure Firewall DNS proxy: diagnose egress before opening rules
  23. 23 Azure Firewall: diagnose a shadowed rule before opening traffic
  24. 24 Azure Firewall IDPS: validate a signature before switching it to deny
  25. 25 Azure route asymmetry: diagnose before changing a UDR
  26. 26 Azure Route Server: validate a BGP advertisement before production propagation
  27. 27 Azure UDR: diagnose an NVA black hole before changing routes
  28. 28 Azure Load Balancer: diagnose a health probe before changing the backend pool
  29. 29 Azure Network Watcher: diagnose an intermittent path before changing NSG or UDR
  30. 30 Azure Network Watcher: validate Flow Logs before opening an NSG rule
  31. 31 Azure Network Watcher: migrate NSG flow logs without losing network evidence
  32. 32 Azure Virtual Network Manager: diagnose a Security Admin rule before changing NSGs
  33. 33 Azure NSG: validate a Service Tag rule before rollout
  34. 34 Azure App Service: validate all-traffic VNet routing before production cutover
  35. 35 Azure Virtual WAN: validate Routing Intent before enforcing production inspection
Automation 4 articles

Automation guardrails

Structure Ansible repositories and AWX job templates so automation stays bounded, reviewable and safe to operate.

Expose useful operations without turning AWX into a remote console.

  1. 01 AWX: design job templates that do not become a dangerous remote console
  2. 02 Ansible in production: structure an operations repository before exposing it in AWX
  3. 03 AWX: rerun a failed job without replaying a partial action
  4. 04 AWX: validate dynamic inventory before a production job
Infrastructure 79 articles

Operations runbooks

Troubleshoot Linux identity, validate Proxmox restores, rotate service identities safely, diagnose Key Vault secret references, qualify Event Grid, Service Bus DLQ and Queue trigger replays, validate Azure DevOps deployment gates, handle Azure Policy denies, bound remediation tasks, rotate APIM subscription keys, and write handover notes that still help after deployment.

Turn architecture and infrastructure into repeatable operations.

  1. 01 Linux and Active Directory: troubleshoot SSSD failures that appear after the join
  2. 02 Proxmox Backup Server: define a testable restore policy, not only a backup policy
  3. 03 Operational architecture documentation: write a note that actually helps run after deployment
  4. 04 Monitoring: turn an alert into an actionable operations runbook
  5. 05 Azure Monitor: diagnose an alert storm after deployment
  6. 06 Azure Monitor and KQL: decide a deployment rollback without silencing alerts
  7. 07 Azure Monitor: diagnose an Action Group before silencing alerts
  8. 08 Service identity and secret rotation: a production runbook, not an isolated task
  9. 09 Azure managed identity: diagnose private access before changing permissions
  10. 10 Azure Managed Identity: diagnose principal drift after resource recreation
  11. 11 Azure Key Vault: migrate access policies to Azure RBAC without breaking workloads
  12. 12 Microsoft Entra PIM: diagnose Azure role activation before granting permanent access
  13. 13 Azure Key Vault: diagnose secret references before rotation
  14. 14 Azure Key Vault: diagnose latency and throttling before rotating secrets
  15. 15 Azure Workload Identity Federation: diagnose CI authentication before bringing back a secret
  16. 16 Azure DevOps WIF: diagnose issuer or subject drift before bringing back a secret
  17. 17 Azure AKS: diagnose Workload Identity before bringing back a secret
  18. 18 Azure Automation: prove the runbook identity before broadening RBAC
  19. 19 Microsoft Entra Workload ID: diagnose Conditional Access before excluding a CI pipeline
  20. 20 Azure Monitor: diagnose missing logs before changing alerts
  21. 21 Azure Monitor: diagnose ingestion latency before changing a KQL alert
  22. 22 Azure Monitor: validate an SLO burn-rate alert before paging on-call
  23. 23 Azure RBAC: diagnose authorization drift before widening a role
  24. 24 Azure Automation: diagnose a runbook before rerunning the job
  25. 25 Azure Functions: diagnose a Timer Trigger before replaying a job
  26. 26 Azure Durable Functions: diagnose an orchestration before replay or purge
  27. 27 Azure Container Apps Jobs: diagnose a KEDA scale rule before rerunning workers
  28. 28 Azure Logic Apps: diagnose a workflow before resubmitting the run
  29. 29 Azure Logic Apps: contain connector throttling before adding retries
  30. 30 Azure Event Grid: diagnose dead-lettered events before replay
  31. 31 Azure Service Bus: diagnose the dead-letter queue before a bounded replay
  32. 32 Azure Functions: diagnose a poison queue before replaying messages
  33. 33 Azure Managed Grafana: diagnose a dashboard or alert before changing KQL
  34. 34 Azure Managed Grafana: diagnose API access before recreating a token
  35. 35 Azure IaC: validate an infrastructure plan before production apply
  36. 36 Terraform: upgrade a provider without uncontrolled drift
  37. 37 Azure Deployment Stacks: validate unmanage and deny settings before an update
  38. 38 Azure Application Gateway: validate a Key Vault certificate before rotation
  39. 39 Azure DevOps: validate approvals and checks before bypassing production
  40. 40 Azure DevOps: diagnose a self-hosted agent before rerunning the pipeline
  41. 41 Azure DevOps: diagnose an offline self-hosted agent before recreating it
  42. 42 Azure DevOps: contain a compromised self-hosted agent before reopening the pool
  43. 43 Azure DevOps: diagnose a queued job before adding self-hosted agents
  44. 44 Azure DevOps: diagnose a Variable Group before rerunning deployment
  45. 45 Azure DevOps: diagnose a canceled deployment before rerunning production
  46. 46 Azure Policy: diagnose a deny before creating a production exemption
  47. 47 Azure Policy: diagnose an expired exemption before disabling the assignment
  48. 48 Azure Policy: bound a remediation task before it rewrites production
  49. 49 Azure APIM: diagnose a subscription key before rotating it in production
  50. 50 Azure Service Health: qualify a regional signal before failover
  51. 51 Azure Traffic Manager: diagnose a degraded profile before forcing failover
  52. 52 Azure Workbooks: diagnose observability drift before changing alerts
  53. 53 Azure Storage: validate lifecycle policy impact before deletion
  54. 54 Azure AKS: unblock a node pool upgrade stopped by a PDB
  55. 55 Azure AKS: diagnose a stuck rollout before forcing rollback
  56. 56 Azure AKS: diagnose HPA oscillation before raising replica limits
  57. 57 Azure Resource Locks: diagnose a blocked deployment before removing the lock
  58. 58 Azure Monitor: roll out a DCR transformation without losing telemetry
  59. 59 Azure Monitor Private Link: diagnose ingestion loss before reopening public access
  60. 60 OpenTelemetry Collector: diagnose telemetry drops before adding memory
  61. 61 OpenTelemetry Collector: validate tail sampling before losing incident traces
  62. 62 Azure Site Recovery: test a recovery plan without touching production
  63. 63 Azure Chaos Studio: bound a resilience test before disrupting production
  64. 64 Azure Resource Graph: reconstruct configuration drift before rollback
  65. 65 Azure Resource Graph: prove inventory completeness before automated remediation
  66. 66 Terraform on Azure: diagnose a partial apply before rerunning production
  67. 67 Terraform on Azure: contain concurrent applies before state divergence
  68. 68 Azure App Configuration: diagnose feature flag drift before rollback
  69. 69 Azure App Service: diagnose TLS renewal before rebinding production
  70. 70 Azure Monitor: diagnose suppressed notifications before changing the alert
  71. 71 Microsoft Sentinel: validate a remediation playbook before production action
  72. 72 Microsoft Sentinel: diagnose a silent scheduled detection before widening its lookback
  73. 73 Azure PostgreSQL: diagnose read replica lag before promotion
  74. 74 Azure Automation: validate a runtime migration before cutting production runbooks over
  75. 75 Azure Backup: validate a VM restore before declaring the recovery path ready
  76. 76 Azure VMSS: diagnose an automatic repair loop before changing the action
  77. 77 Azure SQL: qualify a failover group before a forced failover
  78. 78 Azure Automation: prove the published version before running a runbook
  79. 79 Azure Monitor: contain log alert fan-out before changing the KQL rule
AI 38 articles

Private AI agents

Keep controls around sources, identities, tools, logs and human validation when AI agents operate inside private networks.

Make internal agents useful without making them opaque.

  1. 01 Private-network AI agent: which controls to keep around data, actions and logs
  2. 02 AgentOps: diagnose an AI agent that calls the wrong tool
  3. 03 AgentOps: expose MCP tools without losing control of production actions
  4. 04 AgentOps: diagnose MCP server drift before a production action
  5. 05 Microsoft Foundry: evaluate an agent before giving it a production action
  6. 06 AgentOps: diagnose an AI agent action before rollback
  7. 07 AgentOps: diagnose agent contract drift before redeployment
  8. 08 Azure DevOps MCP: scope an agent before letting it act on the project
  9. 09 AgentOps MCP: validate a tool schema change before redeploying the agent
  10. 10 AgentOps: diagnose missing agent traces before restoring a production action
  11. 11 AgentOps: validate a retrieval index update before it changes production answers
  12. 12 AgentOps: validate an agent tool before it can change production
  13. 13 AgentOps: rotate an agent runtime identity before tool access fails
  14. 14 AgentOps: diagnose Azure OpenAI rate limits before changing models
  15. 15 AgentOps: validate an approval policy before production actions
  16. 16 AgentOps: diagnose prompt injection in retrieval before production actions
  17. 17 Microsoft Foundry: validate guardrails before exposing an agent to production
  18. 18 AgentOps: validate agent memory before production actions
  19. 19 AgentOps: diagnose a failed tool call before retrying a production action
  20. 20 AgentOps: validate the Foundry automation handoff before a production action
  21. 21 AgentOps: validate tool output before a production write
  22. 22 AgentOps: bound MCP tool timeouts, retries and circuit breakers before production
  23. 23 AgentOps: contain an AI agent consumption runaway before shutting down production
  24. 24 AgentOps: shadow a prompt or model change before production cutover
  25. 25 AgentOps: stop a multi-agent handoff loop before actions multiply
  26. 26 Microsoft Foundry Toolbox: validate a version before promoting it to agents
  27. 27 AgentOps: detect evaluation set contamination before promoting an agent
  28. 28 AgentOps: diagnose context-window saturation before raising token limits
  29. 29 AgentOps: revalidate a stale approval before resuming a production action
  30. 30 AgentOps: validate regional failover before rerouting a production agent
  31. 31 AgentOps: invalidate stale evidence before a production action
  32. 32 AgentOps: prevent MCP token forwarding before production actions
  33. 33 Microsoft Foundry: diagnose a content-filter spike before weakening guardrails
  34. 34 AgentOps: calibrate the evaluator before trusting an agent quality gate
  35. 35 AgentOps: neutralize indirect prompt injection in MCP tool output before a production action
  36. 36 AgentOps: validate a fallback model before enabling multi-model routing
  37. 37 AgentOps: contain sensitive data in agent traces before disabling observability
  38. 38 AgentOps: diagnose MCP tool shadowing before a production action