Introduction
In today’s cloud‑first world, the health of your applications, infrastructure, and data pipelines is only as good as the visibility you have into them. Azure Monitor is Microsoft’s unified platform for collecting, analyzing, and acting on telemetry from any Azure resource—virtual machines, containers, databases, IoT devices, and even serverless functions. When you can see metrics and logs in real time, you can spot a failing storage account before a backup slips, scale a web app before traffic spikes, and feed accurate data into self‑governing AI agents that automate remediation.
For Apiary, a platform that tracks bee populations, hive health, and AI‑driven conservation actions, reliable telemetry is a lifeline. Sensors on hives generate temperature, humidity, and acoustic data every few seconds; those streams must be ingested, stored, and correlated with Azure resources that host the analytics dashboards used by researchers worldwide. A solid Azure Monitor setup turns raw numbers into actionable insights—whether you’re preventing a hive overheating or ensuring that an AI agent can autonomously allocate compute for a new machine‑learning model.
This guide walks you through every step required to collect metrics and logs from Azure resources, store them efficiently, query them with Kusto Query Language (KQL), set up alerts, and visualize the data. We’ll also sprinkle in concrete examples, cost considerations, and best‑practice patterns that keep your monitoring stack lean, reliable, and ready for the next big pollination challenge.
1. Understanding Azure Monitor’s Core Components
Azure Monitor isn’t a single service; it’s an ecosystem of tightly integrated pieces. Grasping the architecture early helps you design a monitoring strategy that scales from a single VM to a multi‑region, multi‑tenant deployment.
| Component | What It Does | Typical Use Cases |
|---|---|---|
| Metrics | Numeric, time‑series data (e.g., CPU %, request count). Collected at intervals as low as 1 second for Azure VMs. | Real‑time scaling, performance dashboards. |
| Logs | Structured, semi‑structured, or unstructured records (e.g., Activity Log, Diagnostic Logs). Stored in a Log Analytics workspace. | Auditing, root‑cause analysis, security investigations. |
| Log Analytics Workspace | Central repository for logs, with KQL query engine. Supports retention up to 2 years. | Centralized query across subscriptions, cross‑resource correlation. |
| Azure Monitor Alerts | Rule‑based notifications (email, SMS, webhook, Action Group). Supports metric‑based, log‑based, and smart alerts. | Automated remediation, SLA breach notifications. |
| Workbooks & Dashboards | Interactive visualizations built on top of metrics and logs. | Executive reporting, operational command centers. |
| Azure Monitor for VMs / Containers | Agent‑based collection of guest OS metrics, performance counters, and syslog/event log. | Deep dive into VM health, container performance. |
| Azure Monitor Integration | Connectors to Azure Sentinel, Azure Automation, Logic Apps, and third‑party SIEMs. | Security orchestration, automated workflows. |
Key takeaway: Think of Metrics as the heartbeat, Logs as the medical record, and Workspaces/Alerts as the clinic where you diagnose and treat.
2. Planning Your Data Collection Strategy
Collecting everything sounds safe, but indiscriminate ingestion quickly inflates costs and blurs signal with noise. Follow a data‑first approach:
- Identify Critical Resources – Start with the services that directly affect your bee‑conservation platform:
- Azure Kubernetes Service (AKS) clusters running the analytics pipeline.
- Azure SQL Database storing hive observation data.
- Azure IoT Hub ingesting sensor streams.
- Azure Functions that trigger AI model training.
- Map Required Telemetry – For each resource, list mandatory metrics (e.g.,
CpuPercentage,MemoryRss) and diagnostic logs (e.g.,AppServiceHttpLogs,ContainerLogV2). Use the built‑in Azure Monitor documentation for each service to avoid missing hidden metrics.
- Determine Retention & Granularity –
- Metrics: Default retention is 93 days (1‑minute granularity) and 30 days (1‑second granularity). If you need longer history for trend analysis, enable Metric Retention up to 2 years via the Azure portal or ARM template.
- Logs: Default is 30 days. For compliance (e.g., GDPR data‑processing logs) you may need 365 days. Adjust via the workspace’s Data Retention setting.
- Cost Modeling – Azure Monitor pricing (as of 2024) is:
- Metrics: Free for the first 5 metrics per resource; $0.30 per metric per month thereafter.
- Log Ingestion: $2.30 per GB ingested.
- Log Retention (beyond 30 days): $0.10 per GB per month.
Example: Ingesting 5 GB/day from IoT Hub logs = 150 GB/month → $345 ingestion + $15 retention (if kept 30 days).
- Tag‑Based Governance – Tag resources with
environment:prod|devandmonitoring:enabled. Use Azure Policy azure-policy-monitoring to enforce diagnostic settings on all newly created resources that carry themonitoring:enabledtag.
By defining what, how long, and how often you collect, you avoid the “log‑bloat” trap and keep budgets predictable.
3. Setting Up a Log Analytics Workspace
The Log Analytics workspace is the heart where logs converge. Follow these steps to create a robust workspace that can serve multiple subscriptions and regions.
3.1 Create the Workspace
az monitor log-analytics workspace create \
--resource-group rg-monitoring \
--workspace-name apiary-law \
--location eastus2 \
--sku PerGB2018 \
--retention-time 365
- SKU PerGB2018 gives you pay‑as‑you‑go ingestion pricing, ideal for variable IoT streams.
- Retention‑time 365 stores a full year of logs, sufficient for seasonal bee‑population studies.
3.2 Connect Subscriptions
In the Azure portal, navigate to Workspace > Access Control (IAM) and add the Log Analytics Contributor role to each subscription’s service principal. This enables resources across those subscriptions to write logs without additional configuration.
3.3 Enable Data Sources
Within the workspace UI, select Data sources → Azure Services. Turn on the following connectors (most are toggles):
| Data Source | Recommended Tables | Sample KQL | |
|---|---|---|---|
| Azure Activity Log | AzureActivity | `AzureActivity | summarize count() by Category, bin(TimeGenerated, 1h)` |
| Azure Diagnostics | Perf, InsightsMetrics, AppEvents | `Perf | where ObjectName == "Processor" and CounterName == "% Processor Time"` |
| Azure Security Center | SecurityAlert, SecurityRecommendation | `SecurityAlert | where Severity == "High"` |
3.4 Secure Access
- Read‑Only Role: Assign
Log Analytics Readerto analysts who only need to query. - Data Export: Enable Continuous Export to a storage account for long‑term archiving or to an Event Hub for downstream AI pipelines.
Pro tip: Use Managed Identities for the export process. This eliminates secret rotation and aligns with the principle of least privilege.
4. Configuring Diagnostic Settings for Azure Resources
Diagnostic Settings are the bridge that pushes metrics and logs from a resource into your Log Analytics workspace (or other sinks). Below we walk through three common resource types.
4.1 Azure Storage Account
az monitor diagnostic-settings create \
--resource-id /subscriptions/<sub>/resourceGroups/rg-data/providers/Microsoft.Storage/storageAccounts/hive-logs \
--workspace /subscriptions/<sub>/resourceGroups/rg-monitoring/providers/Microsoft.OperationalInsights/workspaces/apiary-law \
--name storageDiag \
--logs '[{"category":"StorageRead","enabled":true},{"category":"StorageWrite","enabled":true}]' \
--metrics '[{"category":"Transaction","enabled":true,"retentionPolicy":{"days":365,"enabled":true}}]'
- Categories:
StorageRead,StorageWrite,StorageDelete. - Metrics:
Transaction,Ingress,Egress.
These logs allow you to detect anomalous spikes (e.g., a sudden surge in write operations that could indicate a sensor malfunction).
4.2 Azure Kubernetes Service (AKS)
AKS requires two steps: enable Azure Monitor for containers and set up a Log Analytics workspace.
az aks enable-addons \
--resource-group rg-aks \
--name hive-aks \
--addons monitoring \
--workspace-resource-id /subscriptions/<sub>/resourceGroups/rg-monitoring/providers/Microsoft.OperationalInsights/workspaces/apiary-law
Once enabled, the following tables become available: ContainerLog, KubePodInventory, KubeNodeInventory. A typical query to spot OOM kills:
KubePodInventory
| where Reason == "OOMKilled"
| summarize count() by ClusterName, Namespace, bin(TimeGenerated, 1h)
4.3 Azure IoT Hub
IoT Hub diagnostics are crucial for bee‑sensor data pipelines.
az monitor diagnostic-settings create \
--resource-id /subscriptions/<sub>/resourceGroups/rg-iot/providers/Microsoft.Devices/IotHubs/hive-iot \
--workspace /subscriptions/<sub>/resourceGroups/rg-monitoring/providers/Microsoft.OperationalInsights/workspaces/apiary-law \
--name iotHubDiag \
--logs '[{"category":"Connections","enabled":true},{"category":"DeviceTelemetry","enabled":true}]' \
--metrics '[{"category":"D2CMessages","enabled":true}]'
- DeviceTelemetry logs capture each telemetry message (JSON payload).
- Connections logs help you detect rogue devices or connectivity loss.
Tip: Pair the IoT Hub diagnostic logs with Azure Stream Analytics to pre‑filter telemetry before it lands in the workspace, saving up to 40 % on ingestion costs.
5. Working with Metrics Explorer and Alerts
Metrics Explorer is the visual canvas for time‑series data, while Alerts turn those visual cues into automated actions.
5.1 Building a Metrics Chart
- Open Metrics in the Azure portal.
- Select a resource (e.g.,
hive-aks). - Choose Namespace =
Microsoft.ContainerService/managedClusters. - Add Metric =
NodeCPUUtilizationPercentage. - Set Aggregation =
Average. - Apply a Split By on
NodeNameto compare individual nodes.
You can pin this chart to a Dashboard for a real‑time ops view.
5.2 Creating a Smart Alert
Smart alerts use machine learning to detect anomalies beyond static thresholds.
az monitor metrics alert create \
--name "AKS-CPU-Anomaly" \
--resource-group rg-monitoring \
--scopes /subscriptions/<sub>/resourceGroups/rg-aks/providers/Microsoft.ContainerService/managedClusters/hive-aks \
--condition "avg Microsoft.ContainerService/managedClusters/NodeCPUUtilizationPercentage > 80" \
--evaluation-frequency 5m \
--window-size 30m \
--action-group /subscriptions/<sub>/resourceGroups/rg-monitoring/providers/microsoft.insights/actionGroups/ApiaryOps \
--description "Trigger when node CPU stays above 80 % for 30 min"
For a smart alert, replace the condition with DynamicThresholdCriteria. Azure automatically learns the normal range and only fires when the metric deviates significantly.
5.3 Alerting on Log Queries
Log‑based alerts are powered by KQL. Example: Alert when a hive sensor reports temperature > 35 °C for more than 10 minutes.
IoTHubTelemetry
| where DeviceId startswith "hive-"
| where Temperature > 35
| summarize count() by DeviceId, bin(TimeGenerated, 10m)
| where count_ > 0
Create the alert via the portal or CLI:
az monitor scheduled-query create \
--name "Hive-Overheat" \
--resource-group rg-monitoring \
--workspace /subscriptions/<sub>/resourceGroups/rg-monitoring/providers/Microsoft.OperationalInsights/workspaces/apiary-law \
--condition "count > 0" \
--query "IoTHubTelemetry | where DeviceId startswith 'hive-' | where Temperature > 35 | summarize count() by DeviceId, bin(TimeGenerated, 10m)" \
--frequency 5m \
--action /subscriptions/<sub>/resourceGroups/rg-monitoring/providers/microsoft.insights/actionGroups/ApiaryOps
When triggered, the action group can invoke an Azure Function that sends a command to the hive’s cooling system—closing the loop from monitoring to actuation.
6. Visualizing Data with Workbooks and Dashboards
A raw stream of numbers is hard to act on. Workbooks combine KQL queries, charts, and markdown into interactive reports.
6.1 Building a Hive Health Workbook
- In the Log Analytics workspace, select Workbooks → + New.
- Add a Query tile with the following KQL:
IoTHubTelemetry
| where DeviceId startswith "hive-"
| summarize avg(Temperature) by DeviceId, bin(TimeGenerated, 1h)
| render timechart
- Add a second tile for Humidity using a similar query.
- Use Parameters to let users pick a date range or a specific hive ID.
Save the workbook as Hive Health Overview and share it with the research team. Because workbooks are stored as JSON, you can version‑control them in a Git repo alongside your IaC scripts.
6.2 Dashboards for Ops Teams
Dashboards are lightweight and can be pinned to the Azure portal home page.
- Metric Tiles: CPU, Memory, Disk I/O for each AKS node.
- Log Tiles: Count of
SecurityAlertwith severity “High”. - Custom Tile: An embedded Power BI report that visualizes long‑term trends in hive population (pulled from Azure SQL via DirectQuery).
Set Refresh intervals to 1 minute for critical metrics and 5 minutes for logs to balance performance and cost.
7. Automating Responses with Azure Functions and Logic Apps
Monitoring stops being reactive when you automate remediation. Azure Monitor integrates directly with Action Groups, which can invoke Functions, Logic Apps, or even third‑party webhooks.
7.1 Sample Function: Auto‑Scale AI Training Nodes
public static async Task Run(EventGridEvent eventGridEvent, ILogger log)
{
var data = JsonConvert.DeserializeObject<MetricAlertData>(eventGridEvent.Data.ToString());
if (data.CurrentValue > 80) // CPU > 80%
{
var client = new ComputeManagementClient(new DefaultAzureCredential());
await client.VirtualMachineScaleSets.BeginUpdateInstancesAsync(
resourceGroupName: "rg-ml",
vmScaleSetName: "training-set",
instanceIds: new[] { "instanceId" });
log.LogInformation("Scaled out training node due to high CPU.");
}
}
Hook this function to a Metric Alert that watches the TrainingNodeCPU metric. When the alert fires, the function scales the VMSS automatically.
7.2 Logic App for Incident Management
Create a Logic App that:
- Trigger: Azure Monitor Alert (log‑based).
- Action 1: Post a message to a Teams channel (
#apiary-ops). - Action 2: Create a ticket in Azure DevOps with the alert details.
- Action 3: If the alert is a SecurityAlert, invoke an Azure Sentinel playbook that isolates the compromised resource.
The visual designer makes it easy for non‑developers to adjust the workflow—perfect for a conservation organization where staff may rotate frequently.
8. Cost Management and Retention Best Practices
Monitoring is essential, but unchecked ingestion can dwarf your primary compute spend. Follow these proven tactics:
| Practice | How It Works | Savings Estimate |
|---|---|---|
| Sampling | Ingest only 10 % of IoT telemetry using Azure Stream Analytics WHERE RAND() < 0.1. | Up to 90 % reduction in log ingestion. |
| Tiered Retention | Keep 30 days at Hot tier, move older data to Cold (archival storage) via Log Analytics Data Export to a Blob container. | $0.02/GB vs $0.10/GB for hot storage. |
| Metric Aggregation | Use Aggregation (Average, Maximum) at the source instead of storing raw per‑second data. | Reduces metric count; saves $0.30 per metric per month. |
| Diagnostic Settings Scope | Apply diagnostics at resource group level where possible, rather than individually. | Fewer duplicate logs; simplifies policy enforcement. |
| Alert Suppression | Configure Alert Rules with Suppression (e.g., “notify once per hour”). | Cuts down on notification spam and downstream function invocations. |
Use Azure Cost Management + Billing → Cost analysis → Add filter ServiceName eq 'Log Analytics' to monitor real‑time spend. Set a budget alert at 80 % of your monthly monitoring budget to avoid surprises.
9. Integrating Azure Monitor with AI Agents
Self‑governing AI agents thrive on fresh, high‑quality data. Azure Monitor can feed telemetry directly into the training loops of those agents.
9.1 Feeding Metrics into Azure Machine Learning
from azure.monitor.query import MetricsQueryClient
from azure.identity import DefaultAzureCredential
client = MetricsQueryClient(credential=DefaultAzureCredential())
response = client.query_resource(
resource_id="/subscriptions/<sub>/resourceGroups/rg-ml/providers/Microsoft.Compute/virtualMachines/training-node",
metric_names=["Percentage CPU"],
timespan="2024-09-01T00:00:00Z/2024-09-01T01:00:00Z",
granularity="PT1M"
)
cpu_series = response.metrics[0].timeseries[0].data
# Convert to pandas DataFrame and feed to model
The AI agent can predict future load, decide when to spin up additional nodes, or even adjust hyper‑parameters based on current resource utilization.
9.2 Using Log Analytics as a Feature Store
Log tables like IoTHubTelemetry can be queried nightly and persisted into an Azure ML Feature Store:
IoTHubTelemetry
| where TimeGenerated > ago(1d)
| summarize avg(Temperature), max(Humidity) by DeviceId, bin(TimeGenerated, 1h)
| project DeviceId, TimeGenerated, AvgTemp=avg_Temperature, MaxHum=max_Humidity
Export the result to a Parquet file in ADLS Gen2, register it as a feature set, and let downstream models use it for predicting hive health outcomes.
Real‑world example: The Apiary team built an AI agent that automatically recommends supplemental feeding schedules when the model detects a prolonged temperature dip combined with low humidity—using Azure Monitor data as the decision engine.
10. Monitoring for Bee‑Conservation Workloads – A Case Study
To illustrate the concepts, let’s walk through a concrete end‑to‑end scenario.
10.1 Scenario Overview
- Sensors: 150 hives, each streaming temperature, humidity, and acoustic data every 5 seconds to Azure IoT Hub.
- Processing Pipeline: Azure Stream Analytics → Azure Functions (data enrichment) → Azure SQL (historical storage).
- Analytics: Power BI dashboards for researchers; Azure ML models predict colony collapse risk.
10.2 Implementation Highlights
| Step | Azure Service | Configuration |
|---|---|---|
| Telemetry Ingestion | IoT Hub Diagnostic Settings | DeviceTelemetry logs + D2CMessages metric; export to Log Analytics workspace apiary-law. |
| Real‑Time Alert | Metric Alert | D2CMessages > 5000 per minute → webhook to Azure Function that throttles sensor uploads. |
| Log‑Based Alert | Scheduled Query Alert | KQL query detecting Temperature > 35 for > 10 min → Action Group triggers a Logic App that sends a Teams alert and calls the hive’s cooling actuator via Azure IoT Direct Method |