Hosting Workloads on OIT Virtual Infrastructure

Context

 UTD-OIT is running a virtualized cluster at ARDC for virtualizing servers and workloads. 

Benefits for Leveraging OIT’s Virtual Cluster 

Cost Optimization

  • Reduces TCO through consolidation of power, hardware, and support. 
  • Users will only have to manage the application hosted on the VM. 

Performance and Efficiency

  • VMs are hosted on enterprise-grade hardware, fully supported and maintained to ensure peak performance.
  • To maximize uptime, the infrastructure utilizes fully redundant power and networking systems. 

High Availability and Backups

  • The Nutanix cluster was designed to keep applications and VMs running when there are hardware failures. 
  • If a node fails, those VMs will simply be powered on the other nodes in the cluster. 
  • All VMs are backed up for restorability.  

Standards and Guidelines 

OIT is looking to offer this solution to campus partners in an effort to leverage shared resources. This document will cover the standards and guidelines for running these workloads on the platform. 

  • OIT supports and maintains two operating systems- Linux RedHat or Microsoft Windows. 
    • Any potential workload will be required to run only on these two operating systems and abide by the OIT patching and configuration standards for these operating systems. 
  • Console Access will be only given to the OIT System Admins Team, OIT HPC Team, and departmental tech teams. 
    • These groups will help facilitate and maintain the operation of the VM from the operating systems hosted on the VMs. 
  • Application owners and administrators will be given RDP and SSH access to the VMs that their applications are running on. 
    • Level of access to the VM will depend on how application owners and administrators will need to support and maintain their applications. 
  • VM Resource Updates and Changes- You can request more resources like CPU, memory, and storage if needed.
  • .Recovery Points (Snapshots)- This feature is highly recommended to be used when a VM or the application hosted on the VM needs to be updated/upgraded. 
    • Recovery Points are point-in-time copies of a virtual machine’s local disk.  They enable administrators to capture the current state of the VM while allowing it to continue running.  
    • If an upgrade or update fails, you can easily revert the VM to that Recovery Point and restore the VM back to its current state. 
    • Recovery Points can be requested through Atlas by visiting Request Form - Request New VM Snapshot.
  • Backups- Every workload hosted on the OIT Nutanix platform is backed up by Rubrik either once or twice a day. 
    • To find more details and information on OIT’s Backup standards, please visit the following KB (TBD).