Monitoring and Incident Response
This topic details how Communication Recording Agent and Administration Platform systems are monitored at both infrastructure and application levels, and covers incident response procedures.
Quick Highlights
Every aspect of Administration Platform runtime environments is monitored end-to-end.
Use Datadog’s full suite of observation tools, including Infra, Logging, APM and Synthetic Testing
24/7 on call rotation is set up based on service ownership, Security Operations Center for alert monitoring, investigations and incident response for infrastructure and applications.
Automatic alerts (with escalation) are triggered based on specific scenarios or based on anomaly detection.
Incident Response is powered by PagerDuty.
Note
More information can be found at https://www.uniphore.com/security/
Infrastructure Level Monitoring
At the infrastructure level, the following monitoring and protection tools are enabled:
CloudWatch - Monitoring Amazon Web Services (AWS) resources and applications run on AWS in real time.
CloudTrail - Monitoring and Intrusion Detection System (IDS).
Amazon Inspector - Continuous port scanning for vulnerabilities.
GuardDuty - Account and workload threat detection.
VPC Flow logs - Network scanning.
Virus scanning on every compute unit, Wazuh, agent based OS, and server vulnerability scans against known vulnerabilities.
The logs of the above systems are written into a Uniphore Security information and event management (SIEM) and a Network operation center (NOC) monitors the SIEM and looks for patterns.
Note
While it is technically possible to send these logs to an external SIEM if needed, it is generally against Uniphore policy to export these logs externally and it would require exception approval to set this up and export this data outside of Uniphore.
Application Level Monitoring
Uniphore Cloud employs Application Performance Monitoring (APM) solutions from DataDog, across all applications to gain insights into various aspects of their performance. Key functions include:
Error Tracking - APM helps us identify and resolve errors promptly, minimizing service disruptions and ensuring a smooth user experience.
Performance Metrics - By monitoring key performance indicators (KPIs), we can assess the overall health and efficiency of our applications.
Debugging - APM tools aid in identifying and troubleshooting issues, facilitating faster resolution and minimizing downtime.
Monitoring - Continuous monitoring allows us to proactively identify and address performance bottlenecks or anomalies, ensuring optimal application performance at all times.
Uniphore also uses anomaly detection and automatic alerting mechanisms, which play a crucial role in maintaining the reliability and availability of our services. Leveraging advanced monitoring tools, we detect deviations from expected behavior and trigger alerts to notify on-call resources for timely incident handling and service restoration. Key features include:
Datadog Monitors - Monitoring a wide range of conditions affecting running workloads, detecting anomalies and triggering alerts
Anomaly Detection - Some alerts are driven by anomaly detection algorithms, which analyze metrics to identify unusual patterns or deviations from normal behavior.
SLI-based Alerts - The majority of alerts are based on measuring explicit Service Level Indicators (SLIs) against predefined Service Level Objectives (SLOs), ensuring alignment with performance targets and customer expectations.
Communication Recording Agent
Communication Recording Agent leverages the robust APM capabilities of the Uniphore platform and logs and monitors using DataDog. A subset of the service-level monitoring is performed by our CloudOps and DevOps teams including:
Conversations - Captured, Stored, and Processing Time.
Policy - Evaluated, Effective, Ineffective, and Success/Failure.
Service - CPU, Memory, and Up-time.
System Administrator Monitoring
There are 3 types of configurations that the system administrator manages, these services are monitored and reported against:
User and access - The AD system acts as the master system. Any configuration changes here are audited.
Policy Management - All changes done by an administrator are logged in audit logs and are available as APIs as well.
Access Management - If a change is made on User Groups access to data OR functionality, it is also fully logged and audited and available as API or through the UI.
For instructional information on viewing Communication Recording Agent audit logs see Manage Audit Logs.
In addition, the Core Services employ comprehensive audit logging by capturing every action and event within the service. All logs are recorded in an immutable, tamper-proof audit trail, providing full traceability enabling compliance with security and regulatory requirements.