Monday, April 13, 2026

Oracle Database Resiliency Building Blocks and Availability Architecture - PART 1

 

1. What the “Nines” Mean (Availability vs Resiliency)

Availability is usually expressed as:

AvailabilityCommon NameAllowed Downtime / Year
99.9%Three‑nines~8.76 hours
99.99%Four‑nines~52.6 minutes
99.999%Five‑nines~5.26 minutes

👉 Higher nines = less tolerated downtime = much higher architectural complexity and cost


2. Oracle Database Resiliency Building Blocks

Before mapping architectures, these are the Oracle tools used:

  • Oracle Restart – single-node auto-restart
  • Oracle RAC – node-level high availability
  • Oracle Data Guard (DG) – site-level DR (physical standby)
  • Active Data Guard (ADG) – read-only standby + faster failover
  • Fast-Start Failover (FSFO) – automatic DG failover
  • Oracle GoldenGate – logical replication, near-zero data loss
  • Application Continuity / FAN – application resilience
  • Backup & Recovery (RMAN) – last line of defense

3. 99.9% Availability Architecture (Basic HA)

✅ Typical Scenario

  • Internal applications
  • Batch workloads
  • Non-customer-facing systems

🏗️ Oracle Architecture

  • Single Instance Oracle DB
  • Optional:
    • Oracle Restart
    • VM-level HA
  • Backups using RMAN
  • Manual recovery or failover

🔴 Failure Impact

Failure TypeOutcome
DB crashMinutes to hours
OS crashManual restart
Site failureRestore from backup

✅ Summary

  • Low cost
  • Manual intervention
  • Downtime acceptable

4. 99.99% Availability Architecture (Enterprise HA / DR)

✅ Typical Scenario

  • Core enterprise systems
  • ERP, HR, reporting platforms
  • Medium RTO / low RPO

🏗️ Oracle Architecture

Primary Site

  • Oracle RAC (2+ nodes)

DR Site

  • Oracle Data Guard (Physical Standby)
  • Optional Active Data Guard

Automation

  • Data Guard Broker
  • Semi‑automatic failover

🔴 Failure Impact

Failure TypeDowntime
Instance failureSeconds (RAC failover)
Node failureSeconds
DB corruptionMinutes
Site failure5–30 minutes

✅ Summary

  • Zero or near‑zero data loss
  • Fast failover
  • Moderate cost
  • Standard Oracle MAA pattern

5. 99.999% Availability Architecture (Mission‑Critical / Always‑On)

✅ Typical Scenario

  • Banking, trading, telecom
  • Customer-facing 24×7 platforms
  • Regulatory & SLA‑driven systems

🏗️ Oracle Architecture (MAA – Advanced)

Primary Site

  • Oracle RAC (3+ nodes)
  • Enterprise storage with redundancy

Standby Site

  • Active Data Guard with:
    • Fast-Start Failover (FSFO)
    • Observer on third site
  • Or Oracle GoldenGate (for near-zero downtime)

Application Layer

  • Application Continuity
  • FAN / TAF enabled

🔴 Failure Impact

Failure TypeDowntime
Instance failure<5 seconds
Node failure<10 seconds
DB failureAutomatic failover (seconds)
Site failure<1–2 minutes

✅ Summary

  • Automatic failover
  • Near-zero downtime
  • Zero or near-zero data loss
  • High cost & complexity
  • Requires disciplined operations

6. Side‑by‑Side Comparison (Oracle Focused)

Aspect99.9%99.99%99.999%
Oracle RAC❌✅✅
Data Guard❌✅✅
Active Data Guard❌Optional✅
GoldenGate❌❌Optional / ✅
Auto Failover❌Partial✅
Manual OpsHighMediumVery Low
CostLowMediumVery High

7. Key Design Insight (Important)

You don’t achieve five‑nines by just adding technology.
You achieve it by combining:

  • Correct Oracle architecture
  • Application design
  • Network redundancy
  • Storage resilience
  • Well‑tested DR drills
  • Operational maturity

Most outages at 99.999% scale are human or process‑driven, not Oracle failures.

Wednesday, April 8, 2026

Oracle Data Guard 26ai New Features and Enhancements

 Oracle Data Guard 26ai (and 21c onwards) 

  • Faster role transitions – Reduces switchover and failover time by optimizing role change workflows
  • Minimized stall in Maximum Performance – Lowers commit latency while maintaining near‑real‑time protection.
  • Choice of Lag Type for Fast‑Start Failover – Allows failover decisions based on apply lag or transport lag.
  • Faster DML Redirection – Improves performance when redirecting writes during role transitions.
  • Fast‑Start Failover Observer Priority (21c) – Enables prioritized observers to control failover decisions.
  • Up to four Fast‑Start Failover Observers (21c) – Provides higher availability by supporting multiple observers.
  • Rolling Upgrade with Application Continuity – Enables zero‑downtime upgrades while preserving session state.
  • Multiple ASYNC connections – Improves redo transport throughput and resiliency with parallel async paths.
  • Automatic preparation of primary and standby – Automates configuration steps to simplify Data Guard setup.
  • Data Guard Broker PL/SQL API – Allows programmatic management and automation of Data Guard operations.
  • SQLcl support for Data Guard commands (21c) – Enables managing Data Guard directly from SQLcl.
  • ORDS support for Data Guard (21c) – Exposes Data Guard management and monitoring via REST APIs.
  • Show / edit all members at once – Simplifies configuration by allowing bulk updates across members.
  • JSON output for DGMGRL – Provides machine‑readable output for easier integration and automation.
  • Prevent standby databases from becoming primary – Enforces role protection to avoid unintended failovers.
  • Configuration and member tagging – Improves organization and identification of Data Guard components.
  • Automatic standby tempfile creation – Automatically synchronizes tempfiles between primary and standby.
  • PDB Recovery Isolation – Enables recovery operations at the individual PDB level without affecting others.
  • Easy AWR snapshots on the standby – Simplifies performance diagnostics directly on standby databases.
  • Strict database validation – Ensures configuration correctness before switchover or failover operations.
  • Switchover and Failover Readiness – Clearly reports whether the environment is safe for role transitions.
  • Easier tracking of role transitions – Enhances visibility and auditing of role change events.
  • New view V$DG_BROKER_PROPERTY – Provides real‑time visibility into Data Guard Broker properties.
  • New command: VALIDATE DGConnectIdentifier – Verifies network connectivity used by Data Guard.
  • Easier checking of Fast‑Start Failover configurations – Simplifies validation of FSFO setup and health.
  • Fast‑Start Failover Lag Histogram – Visualizes redo lag trends to fine‑tune failover settings.
  • Enhanced observer diagnostic – Improves troubleshooting with richer observer health and status data.
  • Fast‑Start Failover Configuration Validation – Proactively detects configuration issues before failures occur.
  • Offload AI Inference and Vector Search to Oracle Active Data Guard – Runs AI and vector workloads on standby to offload the primary.

Monday, March 23, 2026

Key Differences & 19c Best Practices (AMM and ASMM)

Key Differences between AMM and ASMM & 19c Best Practices


AMM (Automatic Memory Management) -  (MEMORY_TARGET):

How it works: Oracle automatically distributes memory between SGA and PGA.


Pros: Simplest to configure; dynamic resizing across both areas.


Cons: Incompatible with HugePages, which are crucial for performance in large systems. Can have issues with /dev/shm on Linux.


Recommendation: Suitable for testing or small instances (< 4GB SGA).


ASMM (Automatic Shared Memory Management) - (SGA_TARGET):

How it works: Oracle automatically tunes SGA components (buffer cache, shared pool) but PGA is managed separately.


Pros: Recommended for production; supports HugePages for better performance, reduces OS paging.


Cons: Requires manual tuning of PGA_AGGREGATE_TARGET.


19c Context: Oracle 19c generally recommends ASMM (using SGA_TARGET and PGA_AGGREGATE_TARGET) for production environments over AMM to ensure optimal memory management and stability


Complex Oracle 19c to PostgreSQL 15 Database Migration

Production-oriented migration approach for a complex Oracle 19c to PostgreSQL 15 workload , covering assessment, schema and code conversion,...