Difference between revisions of "Sysadmin"

From Earlham CS Department
Jump to navigation Jump to search
(Current Projects)
(Compute (servers and clusters))
 
(191 intermediate revisions by 14 users not shown)
Line 1: Line 1:
__NOTOC__
+
This is the hub for the CS sysadmins on the wiki.
  
= Machines and Brief Descriptions of Services =
+
= Overview =
== CS Machines ==
+
 
[[File:Server_layout_summer2017.jpg|thumb|200px|right|Server layout as of May 2017]]
+
[https://docs.google.com/drawings/d/1XaULz5IxXV_BZQjrko3QJ8wV5aXsSTYcSWxxT49OyZk/edit If you're visually inclined, we have a colorful and easy-to-edit map of our servers here!]
{| style="float:left; margin-right:2px;"
+
 
| style="height:40px; width:150px; text-align:center; background-color:#ADDFFF; border-left:solid 5px #ADDFFF; border-top:solid 5px #ADDFFF; border-bottom:solid 1px white; border-right:solid 5px      #ADDFFF; font-size:120%;" | HOME <br> (vm0)
+
== Server room ==
 +
 
 +
Our servers are in Noyes, the science building that predates the CST. For general information about the server room and how to use it, check out [[Sysadmin:Server Room|this page]].
 +
 
 +
Columns: machine name, IPs, type (virtual, metal), purpose, dies, cores, RAM
 +
 
 +
== Compute Resources ==
 +
 
 +
 
 +
{| class="wikitable"
 +
|+ CS machines and VMs
 +
|-
 +
! Machine name !! 159 Ip Address !! 10Gb Ip address !! Operating System !! Metal or Virtual !! Description !! RAM
 +
|-
 +
| Bowie || 159.28.22.5 || 10.10.10.15 || Debian 9 || Metal || hosts and exports user files; Jupyterhub; landing server || 198 GB
 +
|-
 +
| Smiley || 159.28.22.251 || 10.10.10.252 || Ubuntu 18.04 || Metal || VM host, not accessible to regular users || 156 GB
 +
|-
 +
| Web || 159.28.22.2 || 10.10.10.200 || Ubuntu 18.04 || Virtual || Website host || 8 GB
 +
|-
 +
| Auth || 159.28.22.39 || No 10Gb internet|| CentOS 7 || Virtual || host of LDAP user database || 4 GB
 
|-
 
|-
| style="height:210px; width:150px; background-color:#ADDFFF; border-left:solid 5px #ADDFFF; border-bottom:solid 5px #ADDFFF; border-right:solid 5px #ADDFFF;" | Users <br> SSH <br> NFS <br><br> Backup to Dali: eccs, etc, var
+
| Code || 159.28.22.42 || 10.10.10.42 || Ubuntu 18.04 || Virtual || Gitlab host || 8 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#54C571; border-left:solid 5px #54C571; border-top:solid 5px #54C571; border-bottom:solid 1px white; border-right:solid 5px #54C571; font-size:120%;" | NET <br> (vm1)
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#54C571; border-left:solid 5px #54C571; border-bottom:solid 5px #54C571; border-right:solid 5px #54C571;" | LDAP server <br> [[Sysadmin:DNS & DHCP | DNS]] <br> [[Sysadmin:DNS & DHCP | DHCP]] <br><br> Backup to Dali: etc, var
+
| Net || 159.28.22.1 || 10.10.10.100 || Ubuntu 18.04 || Virtual || network administration host for CS || 4 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#E77471; border-left:solid 5px #E77471; border-top:solid 5px #E77471; border-bottom:solid 1px white; border-right:solid 5px #E77471; font-size:120%;" | WEB <br> (vm2)
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#E77471; border-left:solid 5px #E77471; border-bottom:solid 5px #E77471; border-right:solid 5px #E77471;" | Mailman <br> [[Sysadmin:Mail Stack | Mail Stack]]<br> Apache2 <br> PostgresQL <br> MySQL <br> Wiki <br><br> Backup to Dali: etc, var
+
| Central || 159.28.22.177 || No 10Gb internet || Debian 9 || Virtual || ODK Central Host || 4 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#C38EC7; border-left:solid 5px #C38EC7; border-top:solid 5px #C38EC7; border-bottom:solid 1px white; border-right:solid 5px #C38EC7; font-size:120%;" | TOOLS <br> (vm3)
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#C38EC7; border-left:solid 5px #C38EC7; border-bottom:solid 5px #C38EC7; border-right:solid 5px #C38EC7;" | [[Sysadmin:SageNB Server | SageNB Server]] <br> [[Sysadmin:Jupyterhub Notebook Server | Jupyterhub Server]] <br> [[Sysadmin:Software Modules | Software Modules]] <br> NginX  <br><br> Backup to Dali: etc, var, mnts, sage
+
| Urey || 159.28.22.139 || No 10Gb internet || XCP-ng || Metal || Sysadmin Sandbox Environment || 16 GB
 
|}
 
|}
  
{| style="float:left; margin-right:2px;"
+
{| class="wikitable"
| style="height:55px; width:150px; text-align:center; background-color:#E3A869; border-left:solid 5px #E3A869; border-top:solid 5px #E3A869; border-bottom:solid 1px white; border-right:solid 5px #E3A869; font-size:120%;" | BABBAGE
+
|+ Cluster machines
 +
|-
 +
! Machine name !! 159 Ip Address !! 10Gb Ip address !! Operating System !! Metal or Virtual !! Description !! RAM
 +
|-
 +
| Hopper || 159.28.23.1 || 10.10.10.1 || Debian 10 || Metal || landing server, NFS host for cluster || 64 GB
 +
|-
 +
| Lovelace || 159.28.23.35 || 10.10.10.35 || CentOS 7 || Metal || Large compute server || 96 GB
 +
|-
 +
| Pollock || 159.28.23.8 || 10.10.10.8 || CentOS 7 || Metal || Large compute server || 131 GB
 +
|-  
 +
| Bronte || 159.28.23.140 || No 10Gb internet || CentOS 7 || Metal || Large compute server || 115 GB
 +
|-
 +
| Sakurai || 159.23.23.3 || 10.10.10.3 || Debian 10 || Metal || Runs Backup || 12 GB
 +
|-
 +
| Miyamoto || 159.28.23.45 || No 10Gb currently || Debian 10 || Metal || Runs Backup || 16 GB
 +
|-
 +
| HopperPrime || 159.28.23.142 || 10.10.10.142 || Debian 10 || Metal || Runs Backup || 16 GB
 +
|-
 +
| Monitor || 159.28.23.250 || No 10Gb internet || Debian 11 || Metal || Server Monitoring || 8 GB
 +
|-
 +
| Layout 0 || 159.28.23.2 || 10.10.10.2 || CentOS 7 || Metal || Head Node || 32 GB
 +
|-
 +
| Layout 1 || None || None || CentOS 7 || Metal || Compute Node || 32 GB
 +
|-
 +
| Layout 2 || None || None || CentOS 7 || Metal || Compute Node || 32 GB
 
|-
 
|-
| style="height:210px; width:150px; background-color: #E3A869; border-left:solid 5px #E3A869; border-bottom:solid 5px #E3A869; border-right:solid 5px #E3A869;" | [[Sysadmin:Firewall | Firewall]]
+
| Layout 3 || None || None || CentOS 7 || Metal || Compute Node || 32 GB
|}
+
|-
 
+
| Layout 4 || None || None || CentOS 7 || Metal || Compute Node || 32 GB
{|  
+
|-
| style="height:55px; width:150px; text-align:center; background-color:#EEDC82; border-left:solid 5px #EEDC82; border-top:solid 5px #EEDC82; border-bottom:solid 1px white; border-right:solid 5px #EEDC82; font-size:120%;" | [[Sysadmin:Servers:Proto | PROTO]]
+
| Whedon 0 || 159.28.23.4 || No 10Gb internet|| CentOS 7 || Metal || Head Node || 256 GB
 +
|-
 +
| Whedon 1 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
 +
|-
 +
| Whedon 2 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
 
|-
 
|-
| style="height:210px; width:150px; background-color: #EEDC82; border-left:solid 5px #EEDC82; border-bottom:solid 5px #EEDC82; border-right:solid 5px #EEDC82;" | Weather Monitoring <br> GPS/NTP <br> Energy Monitoring
+
| Whedon 3 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#FF7E6D; border-left:solid 5px #FF7E6D; border-top:solid 5px #FF7E6D; border-bottom:solid 1px white; border-right:solid 5px      #FF7E6D; font-size:120%;" | CONTROL
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#FF7E6D; border-left:solid 5px #FF7E6D; border-bottom:solid 5px #FF7E6D; border-right:solid 5px #FF7E6D;" | Users <br> SSH <br> HOME <br> TOOLS
+
| Whedon 4 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#54C571; border-left:solid 5px #54C571; border-top:solid 5px #54C571; border-bottom:solid 1px white; border-right:solid 5px      #54C571; font-size:120%;" | SMILEY
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#54C571; border-left:solid 5px #54C571; border-bottom:solid 5px #54C571; border-right:solid 5px #54C571;" | [[Sysadmin:XenDocs]] <br> NET <br> WEB
+
| Whedon 5 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#E77471; border-left:solid 5px #E77471; border-top:solid 5px #E77471; border-bottom:solid 1px white; border-right:solid 5px      #E77471; font-size:120%;" | SHINKEN
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#E77471; border-left:solid 5px #E77471; border-bottom:solid 5px #E77471; border-right:solid 5px #E77471;" | Users <br> SSH <br> Add machines
+
| Whedon 6 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:40px; width:150px; text-align:center; background-color:#C38EC7; border-left:solid 5px #C38EC7; border-top:solid 5px #C38EC7; border-bottom:solid 1px white; border-right:solid 5px      #C38EC7; font-size:120%;" |MURPHY
 
 
|-
 
|-
| style="height:210px; width:150px; background-color:#C38EC7; border-left:solid 5px #C38EC7; border-bottom:solid 5px #C38EC7; border-right:solid 5px #C38EC7;" | Elderly email stack <br> Users <br> SSH
+
| Whedon 7 || None || None || CentOS 7 || Metal || Compute Node || 256 GB
|}
 
 
 
<br> <br> <br> <br> <br> <br><br> <br> <br> <br> <br> <br>
 
 
 
== Cluster Machines ==
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px      #0099cc; font-size:120%;" | HOPPER
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc;" | Users <br> SSH <br> NFS server <br> LDAP server <br> [[Sysadmin:Software Modules | Software Modules]] <br> PostgreSQL <br> Wiki <br> Apache2 <br> [[Sysadmin:DNS & DHCP | DNS]] <br> [[Sysadmin:DNS & DHCP | DHCP]]  <br><br> Backup to Dali: etc, var, cluster
+
| Hamilton 0 || 159.28.23.5 || No 10Gb internet || Debian 11 || Metal || Head Node || 128 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#ffdb4d; border-left:solid 5px #ffdb4d; border-top:solid 5px #ffdb4d; border-bottom:solid 1px white; border-right:solid 5px #ffdb4d; font-size:120%;" | DALI
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#ffdb4d; border-left:solid 5px #ffdb4d; border-bottom:solid 5px #ffdb4d; border-right:solid 5px #ffdb4d;" | Storage Server <br>[[Sysadmin:Gitlab | Gitlab]] <br> Backups <br> NginX <br><br> No backup (storage)
+
| Hamilton 1 || None || None || Debian 11 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#ff4d94; border-left:solid 5px #ff4d94; border-top:solid 5px #ff4d94; border-bottom:solid 1px white; border-right:solid 5px #ff4d94; font-size:120%;" | AL-SALAM
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#ff4d94; border-left:solid 5px #ff4d94; border-bottom:solid 5px #ff4d94; border-right:solid 5px #ff4d94;" | WebMO <br> [[Sysadmin:Software Modules | Software Modules]] <br> Apache2 <br><br> No backup
+
| Hamilton 2 || None || None || Debian 11 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#39ad39; border-left:solid 5px #39ad39; border-top:solid 5px #39ad39; border-bottom:solid 1px white; border-right:solid 5px #39ad39; font-size:120%;" | LAYOUT
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#39ad39; border-left:solid 5px #39ad39; border-bottom:solid 5px #39ad39; border-right:solid 5px #39ad39;" | [[Sysadmin:Jupyterhub Notebook Server | Jupyterhub Server]] <br> [[Sysadmin:Software Modules | Software Modules]] <br> NginX <br> Apache2 <br> WebMO <br><br> Backup to Dali: etc, var
+
| Hamilton 3 || None || None || Debian 11 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px #0099cc; font-size:120%;" | BRONTE
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc;" | [[Sysadmin:Software Modules | Software Modules]] <br><br> Backup to Dali: etc, var, nbserver
+
| Hamilton 4 || None || None || Debian 11 || Metal || Compute Node || 256 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px #0099cc; font-size:120%;" | POLLOCK
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc;" | [[Sysadmin:Software Modules | Software Modules]] <br> WebMO <br> NginX <br><br> No backup
+
| Hamilton 5 || None || None || Debian 11 || Metal || Compute Node || 256 GB
 
|}
 
|}
  
{| style="float:left; margin-right:2px;"
+
{| class="wikitable"
| style="height:55px; width:150px; text-align:center; background-color:#ffdb4d; border-left:solid 5px #ffdb4d; border-top:solid 5px #ffdb4d; border-bottom:solid 1px white; border-right:solid 5px #ffdb4d; font-size:120%;" | KAHLO
+
|+ Lab machines
 
|-
 
|-
| style="height:300px; width:150px; background-color:#ffdb4d; border-left:solid 5px #ffdb4d; border-bottom:solid 5px #ffdb4d; border-right:solid 5px #ffdb4d;" | Storage Server <br>Backups <br> NginX <br><br> No backup
+
! Machine name !! 159 Ip Address !! Location !! Operating System !! RAM
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px #0099cc; font-size:120%;" | BIGFE
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc;" | [[Sysadmin:Software Modules | Software Modules]]
+
| Borg || 159.28.22.10 || Turing (CST 222) || Ubuntu 20 || 16 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px #0099cc; font-size:120%;" | T-VOC
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc;" | [[Sysadmin:Software Modules | Software Modules]]
+
| Gao || 159.28.22.11 || Turing (CST 222) || Ubuntu 20 || 8 GB
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:150px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px #0099cc; font-size:120%;" | ELWOOD
 
 
|-
 
|-
| style="height:300px; width:150px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc;" | [[Sysadmin:Software Modules | Software Modules]]
+
| Snyder || 159.28.22.12 || Turing (CST 222) || Ubuntu 20 || 8 GB
|}
 
 
 
 
 
 
 
 
 
 
 
 
 
<br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br>
 
 
 
== Switches ==
 
 
 
 
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:175px; text-align:center; background-color:#0099cc; border-left:solid 5px #0099cc; border-top:solid 5px #0099cc; border-bottom:solid 1px white; border-right:solid 5px      #0099cc; font-size:120%;" | SG538SF02J
 
 
|-
 
|-
| style="height:200px; width:175px; background-color:#0099cc; border-left:solid 5px #0099cc; border-bottom:solid 5px #0099cc; border-right:solid 5px #0099cc; font-size:80%;" |  
+
| Goldwasser || 159.28.22.13 || Lovelace (CST 219) || Ubuntu 20 || 8 GB
*Model: HP Procurve 3400cl
 
*Ports: 24
 
*Backplane bandwidth:
 
**88 Gbps
 
**64 million pps
 
*Memory:
 
**2MB packet buffer
 
**16 MB dual flash
 
**128 MB SDRAM
 
*Cut-through switching: No
 
*Unused as of May 12, 2017
 
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:175px; text-align:center; background-color:#ffdb4d; border-left:solid 5px #ffdb4d; border-top:solid 5px #ffdb4d; border-bottom:solid 1px white; border-right:solid 5px #ffdb4d; font-size:120%;" | CN63FP762S
 
 
|-
 
|-
| style="height:200px; width:175px; background-color:#ffdb4d; border-left:solid 5px #ffdb4d; border-bottom:solid 5px #ffdb4d; border-right:solid 5px #ffdb4d;font-size:80%;" |  
+
| Bartik || 159.28.22.14 || Lovelace (CST 219) || Ubuntu 20 || 8 GB
*Model: HP 2530-24G
 
*Ports: 24
 
*Switching Capacity:
 
**56 Gbps
 
**41.6 million pps
 
*Memory:
 
**1.5 MB packet buffer
 
**256 MB  flash
 
**128 MB DDR3 DIMM
 
*Cut-through switching: No
 
*Connected to Al-Salam as of May 12, 2017
 
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:175px; text-align:center; background-color:#ff4d94; border-left:solid 5px #ff4d94; border-top:solid 5px #ff4d94; border-bottom:solid 1px white; border-right:solid 5px #ff4d94; font-size:120%;" | SG525SG025
 
 
|-
 
|-
| style="height:200px; width:175px; background-color:#ff4d94; border-left:solid 5px #ff4d94; border-bottom:solid 5px #ff4d94; border-right:solid 5px #ff4d94;font-size:80%;" |  
+
| Wilson || 159.28.22.15 || Lovelace (CST 219) || Ubuntu 20 || 8 GB
*Model: HP Procurve 3400cl
 
*Ports: 24
 
*Backplane bandwidth:
 
**88 Gbps
 
**64 million pps
 
*Memory:
 
**2MB packet buffer
 
**16 MB dual flash
 
**128 MB SDRAM
 
*Cut-through switching: No
 
*Connected to layout and whedon as of May 12, 2017
 
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:175px; text-align:center; background-color:#39ad39; border-left:solid 5px #39ad39; border-top:solid 5px #39ad39; border-bottom:solid 1px white; border-right:solid 5px #39ad39; font-size:120%;" | Netgear JGS524
 
 
|-
 
|-
| style="height:200px; width:175px; background-color:#39ad39; border-left:solid 5px #39ad39; border-bottom:solid 5px #39ad39; border-right:solid 5px #39ad39;font-size:80%;" |  
+
| Bilas || 159.28.22.16 || Lovelace (CST 219) || Ubuntu 20 || 8 GB
*Current cluster head-node
 
*Unmanaged (no console/configuration)  
 
*Ports: 24
 
*Switching bandwidth:
 
**48 Gbps
 
**1.5 million pps
 
*Memory:
 
**2MB packet buffer
 
*Cut-through switching: No
 
*Connected to Al-Salam, Hopper, Pollock, Nagios, Dali, Kahlo, Bronte as of May 12, 2017
 
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:175px; text-align:center; background-color:#E77471; border-left:solid 5px #E77471; border-top:solid 5px #E77471; border-bottom:solid 1px white; border-right:solid 5px #E77471; font-size:120%;" | cs-main
 
 
|-
 
|-
| style="height:200px; width:175px; background-color:#E77471; border-left:solid 5px #E77471; border-bottom:solid 5px #E77471; border-right:solid 5px #E77471;font-size:80%;" |  
+
| Johnson || 159.28.22.17 || Lovelace (CST 219) || Ubuntu 20 || 8 GB
*Model: HP 5920AF-24XG
 
*Ports: 24
 
*Backplane bandwidth:
 
**480 Gbps
 
**367 million pps
 
*Memory:
 
**3.6 GB packet buffer
 
**256 MB dual flash
 
**2 GB SDRAM
 
*Cut-through switching: Yes
 
*IP Address: 159.28.31.66
 
*Connected to layout, kahlo, and dali as of May 12, 2017
 
|}
 
 
 
{| style="float:left; margin-right:2px;"
 
| style="height:55px; width:175px; text-align:center; background-color:#ADDFFF; border-left:solid 5px #ADDFFF; border-top:solid 5px #ADDFFF; border-bottom:solid 1px white; border-right:solid 5px #ADDFFF; font-size:120%;" | 5500denniscs-sw1
 
 
|-
 
|-
| style="height:200px; width:175px; background-color:#ADDFFF; border-left:solid 5px #ADDFFF; border-bottom:solid 5px #ADDFFF; border-right:solid 5px #ADDFFF;font-size:80%;" |  
+
| Graham || 159.28.22.14 || Lovelace (CST 219) || Ubuntu 20 || 8 GB
*Model: HP 5500 JG542A
 
*Ports: 24
 
*Backplane bandwidth:
 
**224 Gbps
 
**166.6 million pps
 
*Memory:
 
**6 MB packet buffer
 
**512 MB dual flash
 
**1 GB SDRAM
 
*Cut-through switching: No
 
*IP Address: 159.28.31.67
 
*Connected to Babbage, Control, Nagios, and the cluster's netgear switch (via port 14) as of May 12, 2017
 
 
|}
 
|}
  
 +
=== CS Machine Address List ===
 +
<pre>bowie.cs.earlham.edu smiley.cs.earlham.edu web.cs.earlham.edu auth.cs.earlham.edu code.cs.earlham.edu net.cs.earlham.edu central.cs.earlham.edu urey.cs.earlham.edu</pre>
  
 +
=== Cluster Machine Address List ===
 +
<pre>
 +
hopper.cluster.earlham.edu lovelace.cluster.earlham.edu pollock.cluster.earlham.edu bronte.cluster.earlham.edu sakurai.cluster.earlham.edu miyamoto.cluster.earlham.edu hopperprime.cluster.earlham.edu monitor.cluster.earlham.edu whedon.cluster.earlham.edu layout.cluster.earlham.edu hamilton.cluster.earlham.edu</pre>
  
 +
=== Lab Machine Address List ===
 +
<pre>borg.cs.earlham.edu gao.cs.earlham.edu snyder.cs.earlham.edu goldwasser.cs.earlham.edu bartik.cs.earlham.edu wilson.cs.earlham.edu bilas.cs.earlham.edu johnson.cs.earlham.edu graham.cs.earlham.edu</pre>
  
 +
=== Specialized resources ===
  
 +
Specialized computing applications are supported on the following machines:
  
 +
* [[Sysadmin:GPGPU|GPU’s for AI/ML/data science]]: layout cluster
 +
* virtualization: smiley
 +
* containers: bowie
  
 +
== Network ==
  
 +
We have two network fabrics linking the machines together. There are three subdomains.
  
 +
=== 10 Gb ===
  
 +
We have 10Gb fabric to mount files over NFS. Machines with 10Gb support have an IP address in the class C range 10.10.10.0/24 and we want to add DNS to these addresses.
  
 +
=== 1 Gb (cluster, cs) ===
  
 +
We have two class C subnets on the 1Gb fabric: 159.28.22.0/24 (CS) and 159.28.23.0/24 (cluster). This means we have double the IP addresses on the 1Gb fabric that we have on the 10Gb fabric.
  
 +
Any user accessing *.cluster.earlham.edu and *.cs.earlham.edu is making calls on a 1Gb network.
  
<br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br><br>
+
=== Intra-cluster fabrics ===
  
= Systems Administration Documentation =
+
The layout cluster has an Infiniband infrastructure. Wachowski has only a 1Gb infrastructure.
For old documentation, see: [[Sysadmin:Old | Old Wiki Information]]
 
  
{|
+
== Power ==
|- valign:"top"
 
|
 
<div style="border:10px solid #E0EAF8; padding:5px; width:230px; height:500px">
 
<div style="background-color:#CEDEF4; padding:5px;">
 
 
 
=== Admin Tasks ===
 
</div>
 
* [[Sysadmin:Nagios | Nagios Monitoring ]]
 
* [[Sysadmin:Shinken | Shinken Monitoring ]]
 
* [[Sysadmin:Upgrading SSL Certificate | Upgrading SSL Certificates ]]
 
* [[Sysadmin:User Management | User Management]]
 
* [[Newmodules | Installing software under modules ]]
 
* [[Sysadmin:Backup|Backup]]
 
* [[Sysadmin:Contacting all users|Contacting all users]]
 
* [[Sysadmin:New Sysadmins | Welcoming a new sysadmin to the fold]]
 
* [[Sysadmin:AddComputer|Add a computer]]
 
* [[Sysadmin:Setting up Lovelace Lab Machines | Setting up Lovelace Lab Machines]]
 
* [[Reset password]]
 
  
 +
We have a backup power supply, with batteries last upgraded in 2019 (?). We’ve had a few outages since then and power has held up well.
  
<!-- This has to stay as part of the formatting -->
+
== HVAC ==
</div>
 
| style="float:left;" |
 
|
 
<div style="border:10px solid #FFDFFF; padding:5px; width:230px; height:500px;">
 
<div style="background-color:#FFCEFF; padding:5px;">
 
  
=== Services ===
+
HVAC systems are static and are largely managed by Facilities.
</div>
 
* [[Sysadmin:Services:ClusterOverview|Cluster Overview]]
 
* [[Sysadmin:Services:Apache2|Apache2]]
 
* [[Sysadmin:Services:Databases|Databases]]
 
* [[Sysadmin:DNS & DHCP|DNS and DHCP]]
 
* [[Sysadmin:Services:Virtualization | Virtualization]]
 
* [[Sysadmin:Services:XenServerSetup | Xen Server]]
 
  
<!-- This has to stay as part of the formatting -->
+
[[Topology|See full topology diagrams here.]]
</div>
 
| style="float:left;" |
 
|
 
<div style="border:10px solid #F0DDD5; padding:5px; width:230px; height:500px;">
 
<div style="background-color:#E4C0B1; padding:5px;">
 
  
=== Miscellaneous ===
+
[[Sysadmin:Layers of abstraction for filesystems|A word about what's happening between files and the drives they live on.]]
</div>
 
* [[ShutdownProcedure| Shutdown and Boot up]]
 
* [[SysadminContactInfo| Contact Information]]
 
* [[Sysadmin:ImportantInfo:PhoneNumbers| Phone Numbers]]
 
* [[Sysadmin:ImportantInfo:AuthenticationInfo| Authentication Information]]
 
* [[Sysadmin:ImportantInfo:UPS| UPS]]
 
* [[Sysadmin:ImportantInfo:SSLcerts| Generating SSL Certificates]]
 
* [[Sysadmin:Power draws| Power draws]]
 
* [[Sysadmin:ImportantInfo:SunHardware|Working with Sun Hardware]]
 
* [[Sysadmin:Passwords]]
 
* Patching
 
** [[LinuxKernelPatching|Linux Kernel Patching]]
 
* [[Sysadmin:SerialConsoleCableEnds|Cable Ends]]
 
* [[Sysadmin:VirtualizationComparison|NEW Virtualization Comparison]]
 
  
<!-- This has to stay as part of the formatting -->
 
</div>
 
| style="float:left;" |
 
|
 
<div style="border:10px solid #D6F8DE; padding:5px; width:230px; height:500px;">
 
<div style="background-color:#BDF4CB; padding:5px;">
 
  
=== Networking ===
+
= New sysadmins =
</div>
 
* [[Sysadmin:Networking:NetworkLayout|Network Layout (as of 08/2006)]]
 
* [[Sysadmin:Networking:D224 cable plant|D224 cable plant]]
 
* [[Sysadmin:Networking:Fiber plans|Fiber plans]]
 
* [[Sysadmin:Networking:Switches|Switches]]
 
* [[Sysadmin:Networking:Rack notes|Rack notes]]
 
* [[Sysadmin:Networking:Public|Public Network]]
 
* [[Sysadmin:Networking:NetworkTopo|Old Network Topo Figures]]
 
* [[Sysadmin:Networking:NetworkDiagram|Network layout (May 2007)]]
 
* [[Sysadmin:Networking:Alternate Network Path|Alt Network path]]
 
* [[Sysadmin:UPS Setup]]
 
  
<!-- This has to stay as part of the formatting -->
+
These pages will be helpful for you if you're just starting in the group:
</div>
 
|}
 
  
== Current Projects ==
+
* [[Sysadmin:New Sysadmins | Welcoming a new sysadmin ]]
=== Last updated 22 March 2018 ===
+
* [[Sysadmin:Troubleshooting|General troubleshooting tips for admins]]
* lovelace machine to ITS - Eli
+
* [[Sandbox Notes|Sandbox Notes]]
* new folks - have machines, have networking, ready to do the installs - Alek
+
* [[Password managers]]
* 10 Gbps for dali and kahlo, hopper, etc (nfs mounts in /etc/fstab) - Adam
+
* [[Server safety]]
* Layout - head node swap, check disk space on /scratch - Adam
+
* [https://code.cs.earlham.edu/sysadmin/ticket-tracker Ticket tracking for current projects]
* Shinken - compute nodes, APC UPS - Alek
 
* Backup - max disk capacity, scripts on all machines backing-up at least /etc; babbage - Chau
 
* Ganglia on Hopper - Ahsan
 
* UPS load monitoring - Eli
 
  
* -------------------------
+
Note: you'll need to log in with wiki credentials to see most Sysadmin pages.
* FIFO for requests rather than ad-hoc
 
* Accounting for hours logged
 
* PBS Shinken monitoring
 
* Power layout - legend, color-code servers by type, how are servers with 2x power supplies plumbed?
 
  
* -------------------------
+
= Additional information =
* Shinken - Vitalii and Aleks (documentation, monitoring webmo and pm8)
 
* Hadoop on Whedon - Vitalii and Adam (stuck on ?)
 
* Layout - Adam (stuck on ldap)
 
* Gaussian & WebMO on Whedon - Ahsan and Eli (stuck on firewall)
 
* Backup - Ch'''â'''u (moving along, setup backup.cs.e.e next)
 
* Installing power monitor, etc. and rack cleanup - TO BE ASSIGNED (eli and charlie)
 
** switch switch (charlie to check inventory)
 
* Mothur - Ahsan
 
* password policy, force change and random initial
 
** for now notify people with default and then change after a couple of days; script will generate random string
 
* Talk about at next meeting:
 
** Spring break and summer people (important)
 
** Jon's user and Postgres database
 
** investigate tools /clients/ directory with what looks like duplicate user directories
 
  
=== (list from 2017-10-26) ===
+
These pages contain a lot of the most important information about our systems and how we operate.
* <s>Finish migrating tools and home to smiley</s> migrate web and net back to control
 
** Record consistent & thorough documentation, especially concerning the startup and shutdown of the VMs
 
* Setup graceful shutdown when we detect to be running solely off UPS
 
** Additionally, setup clean shutdown and startup for VMs on <s>smiley</s> control (?)
 
* Fix reverse lookup error for mail.cs.earlham.edu
 
** Should consistently refer to 159.28.22.2 (web.cs.earlham.edu)
 
** It's possible that this isn't actually broken.
 
* Layout infiniband subnet manager
 
* Layout disk swap, new lo0
 
** <s>Redo /scratch for mglerner group on /media/r10_vol</s> ?
 
* Migrate Elwood, BigFe, t-voc to repurposed Lovelace Machines (Eli)
 
* <s>HP Al-Salam switch enable jumboframes</s> ?
 
* Strike unused lovelace machine addresses from CS DNS file
 
** Perhaps there's a python file in root's home somewhere that checks for unused DNS/DHCP addresses?
 
  
== Ongoing Projects (Spring 2017) ==  
+
===Technical docs===
=== TODO ===
 
* EMAILING ALL THE USERS https://wiki.cs.earlham.edu/index.php/Sysadmin:Old:Contacting_All_Users
 
* SHUTDOWN SCHEDULED FOR SUNDAY (APRIL 16)
 
** Check/update instructions - one version is at https://wiki.cs.earlham.edu/index.php/Sysadmin:ImportantInfo:PowerFailure, there are others too
 
** Notify users
 
* Fix certs for gitlab, etc.
 
* Secure 1-2 admins for the summer
 
* Prep layout for May-June usage
 
* Practice shutdown-startup procedure (with Michael)
 
* Nsswitch consistency across all machines
 
* Document tools: startup / shutdown - Charlie
 
* Use Sysadmin namespace for all our pages - All
 
** Testing usefulness of documentation - Dave
 
* Al Salam: configure switch, re-rack. - Vitalii
 
** HP switch should be reset and tested.
 
* LDAP cleanup of system users / old groups - James
 
* Layout - Nirdesh
 
** Lo0 RAID (mdadm)
 
** 10GB from Dali to lo0 (adding rules on compute node routing tables as a possible fix)
 
** BIOS reset
 
* 10Gb, perfsonar, ...
 
* Monitoring: (Ganglia, Shinken)
 
** Getting consistency among all the machines(check_nrpe regularly stops working).
 
* Whedon: configured and available
 
* Change passwords (on everything). Postgres, shenken, ...
 
* Webcam on office whiteboard (new office location?)
 
* Learn virtual machine architecture and modules - Dave
 
** Document in a format for future admin training?
 
** Find existing introduction material
 
* Mirror ''control'' for testing, swapping, etc.
 
  
=== DONE (19 Jan 2017) ===
+
* [https://code.cs.earlham.edu/sysadmin/ticket-tracker Ticket tracking for current projects]
* Examine extra "layout" node. - Adam
+
* [[Server safety]]
** Differences are: Single PSU, Single GPGPU, No VGA.
+
* [[Sysadmin:Backup|Backup]]
** It has Infiniband and 10GB cards installed.
+
* [[Sysadmin:Monitoring | Monitoring ]]
* Networking - Adam, Charlie
+
* [[Sysadmin:SSH|SSH info relevant to admins]]
** IP over Infiniband working on layout
+
* [[Sysadmin:User Management | User Management]] and [[Sysadmin:LDAP|LDAP]] generally
*** Resolved by resetting IB switch configuration: <code>ibwarn: [3349] mad_rpc_rmpp: _do_madrpc failed; dport (Lid 1)</code>
+
* [[Sysadmin:Jupyterhub Notebook Server|Jupyterhub]] and [[Nbgrader notes|NBGrader]]
 +
* [[Sysadmin:MailStack|Email service]]
 +
* [[Sysadmin:XenDocs | Xen Server]]
 +
* [[Sysadmin:NFS|Network File System (NFS)]]
 +
* [[Sysadmin:Web Servers|Web Servers and Websites]]
 +
* [[Sysadmin:Services:Databases|Databases]]
 +
* [[Sysadmin:DNS & DHCP|DNS and DHCP]]
 +
* [[Sysadmin:AWS|AWS]]
 +
* [[Bash_start_up_script|Bash startup scripts]]
 +
* [[Sysadmin:VirtualBox | VirtualBox]]
 +
* [[X Applications]]
 +
* [[Sysadmin:Services:ClusterOverview|Cluster Overview]] and [[Sysadmin:Ccg-admin|additional details]]
 +
* [[Sysadmin:Firewall|Firewall]] running on babbage.cs.e.e
 +
* [[Sysadmin:Setting_up_Lovelace_Lab_Machines|Setting up Lab Machines]]
  
=== FUTURE ===
+
===Common tasks===
* Centralized password database / manager / location
+
* [[Sysadmin:Recurring Tasks | Recurring tasks - e.g. software updates, hardware replacements]]
 +
* [[Sysadmin:Contacting all users|Contacting all users]]
 +
* [[Reset password]]
 +
* [[Sysadmin:Software installation | Software installation]]
 +
* [[Modules | Installing software under modules ]]
 +
* [[Sysadmin:AddComputer|Add a computer to CS or cluster domains]]
 +
* [[Senior projects|Supporting senior projects]]
 +
* [[ShutdownProcedure|How to do a planned shutdown and reboot of the system]]
 +
** [[Sysadmin:TestingServices | Testing services]] (after a reboot, upgrade, change in the phase of the moon, etc.)
 +
* [[Sysadmin:Upgrading SSL Certificate | Upgrading SSL Certificates ]]
 +
* [[Sysadmin:Launch at startup|Launch a process at startup]]
 +
* [[Sysadmin:Psql-setup | setup psql for cs430 students]]
  
== Current Projects (updated 13 Oct 16) ==  
+
===Group and institution information===
* '''Groups and LDAP and sudo - James'''
+
* [[Sysadmin:CS-ITS Interoperability|Working with ITS]]
* <s>Amber - James</s>
+
* [[Sysadmin:Recurring spending | Recurring spending ]]
* <s>Edward's setup - Vitalli</s>
+
* [[Sysadmin:SlackAndGitLab | Slack and GitLab integration]]
* <s>WebDev access - Nirdesh<s>
 
* Puppet - James and Vitalii
 
* '''Bacula - Nirdesh'''
 
* SSL certificate upgrade and documentation - Kristin
 
* <s>Listserv merging with archives preserved - Nirdesh </s>
 
* '''Ganglia - Bret'''
 
* '''Shenken - Vitalii'''
 
** latency, UPS
 
* New Layout node - ? and ?
 
* Provision Sappho (compute) - after Puppet
 
* Provision Kahlo (storage) -
 
** replace broken drive
 
* I2 setup
 
** DTN, storage nodes, head nodes, ports in CST
 
* [[Sysadmin:WhedonProvisioning|Provision Whedon]] (compute) - after Puppet
 
* '''Shutdown and startup test - scheduled for Sunday 27 November'''
 
* Disk cleaning - Charlie
 
* <s>Password changing in the CS and cluster domains - Vitalii and James</s>
 
* Proto setup and maintenance with HIP/Green Science
 

Latest revision as of 11:15, 1 June 2022

This is the hub for the CS sysadmins on the wiki.

Overview

If you're visually inclined, we have a colorful and easy-to-edit map of our servers here!

Server room

Our servers are in Noyes, the science building that predates the CST. For general information about the server room and how to use it, check out this page.

Columns: machine name, IPs, type (virtual, metal), purpose, dies, cores, RAM

Compute Resources

CS machines and VMs
Machine name 159 Ip Address 10Gb Ip address Operating System Metal or Virtual Description RAM
Bowie 159.28.22.5 10.10.10.15 Debian 9 Metal hosts and exports user files; Jupyterhub; landing server 198 GB
Smiley 159.28.22.251 10.10.10.252 Ubuntu 18.04 Metal VM host, not accessible to regular users 156 GB
Web 159.28.22.2 10.10.10.200 Ubuntu 18.04 Virtual Website host 8 GB
Auth 159.28.22.39 No 10Gb internet CentOS 7 Virtual host of LDAP user database 4 GB
Code 159.28.22.42 10.10.10.42 Ubuntu 18.04 Virtual Gitlab host 8 GB
Net 159.28.22.1 10.10.10.100 Ubuntu 18.04 Virtual network administration host for CS 4 GB
Central 159.28.22.177 No 10Gb internet Debian 9 Virtual ODK Central Host 4 GB
Urey 159.28.22.139 No 10Gb internet XCP-ng Metal Sysadmin Sandbox Environment 16 GB
Cluster machines
Machine name 159 Ip Address 10Gb Ip address Operating System Metal or Virtual Description RAM
Hopper 159.28.23.1 10.10.10.1 Debian 10 Metal landing server, NFS host for cluster 64 GB
Lovelace 159.28.23.35 10.10.10.35 CentOS 7 Metal Large compute server 96 GB
Pollock 159.28.23.8 10.10.10.8 CentOS 7 Metal Large compute server 131 GB
Bronte 159.28.23.140 No 10Gb internet CentOS 7 Metal Large compute server 115 GB
Sakurai 159.23.23.3 10.10.10.3 Debian 10 Metal Runs Backup 12 GB
Miyamoto 159.28.23.45 No 10Gb currently Debian 10 Metal Runs Backup 16 GB
HopperPrime 159.28.23.142 10.10.10.142 Debian 10 Metal Runs Backup 16 GB
Monitor 159.28.23.250 No 10Gb internet Debian 11 Metal Server Monitoring 8 GB
Layout 0 159.28.23.2 10.10.10.2 CentOS 7 Metal Head Node 32 GB
Layout 1 None None CentOS 7 Metal Compute Node 32 GB
Layout 2 None None CentOS 7 Metal Compute Node 32 GB
Layout 3 None None CentOS 7 Metal Compute Node 32 GB
Layout 4 None None CentOS 7 Metal Compute Node 32 GB
Whedon 0 159.28.23.4 No 10Gb internet CentOS 7 Metal Head Node 256 GB
Whedon 1 None None CentOS 7 Metal Compute Node 256 GB
Whedon 2 None None CentOS 7 Metal Compute Node 256 GB
Whedon 3 None None CentOS 7 Metal Compute Node 256 GB
Whedon 4 None None CentOS 7 Metal Compute Node 256 GB
Whedon 5 None None CentOS 7 Metal Compute Node 256 GB
Whedon 6 None None CentOS 7 Metal Compute Node 256 GB
Whedon 7 None None CentOS 7 Metal Compute Node 256 GB
Hamilton 0 159.28.23.5 No 10Gb internet Debian 11 Metal Head Node 128 GB
Hamilton 1 None None Debian 11 Metal Compute Node 256 GB
Hamilton 2 None None Debian 11 Metal Compute Node 256 GB
Hamilton 3 None None Debian 11 Metal Compute Node 256 GB
Hamilton 4 None None Debian 11 Metal Compute Node 256 GB
Hamilton 5 None None Debian 11 Metal Compute Node 256 GB
Lab machines
Machine name 159 Ip Address Location Operating System RAM
Borg 159.28.22.10 Turing (CST 222) Ubuntu 20 16 GB
Gao 159.28.22.11 Turing (CST 222) Ubuntu 20 8 GB
Snyder 159.28.22.12 Turing (CST 222) Ubuntu 20 8 GB
Goldwasser 159.28.22.13 Lovelace (CST 219) Ubuntu 20 8 GB
Bartik 159.28.22.14 Lovelace (CST 219) Ubuntu 20 8 GB
Wilson 159.28.22.15 Lovelace (CST 219) Ubuntu 20 8 GB
Bilas 159.28.22.16 Lovelace (CST 219) Ubuntu 20 8 GB
Johnson 159.28.22.17 Lovelace (CST 219) Ubuntu 20 8 GB
Graham 159.28.22.14 Lovelace (CST 219) Ubuntu 20 8 GB

CS Machine Address List

bowie.cs.earlham.edu smiley.cs.earlham.edu web.cs.earlham.edu auth.cs.earlham.edu code.cs.earlham.edu net.cs.earlham.edu central.cs.earlham.edu urey.cs.earlham.edu

Cluster Machine Address List

hopper.cluster.earlham.edu lovelace.cluster.earlham.edu pollock.cluster.earlham.edu bronte.cluster.earlham.edu sakurai.cluster.earlham.edu miyamoto.cluster.earlham.edu hopperprime.cluster.earlham.edu monitor.cluster.earlham.edu whedon.cluster.earlham.edu layout.cluster.earlham.edu hamilton.cluster.earlham.edu

Lab Machine Address List

borg.cs.earlham.edu gao.cs.earlham.edu snyder.cs.earlham.edu goldwasser.cs.earlham.edu bartik.cs.earlham.edu wilson.cs.earlham.edu bilas.cs.earlham.edu johnson.cs.earlham.edu graham.cs.earlham.edu

Specialized resources

Specialized computing applications are supported on the following machines:

Network

We have two network fabrics linking the machines together. There are three subdomains.

10 Gb

We have 10Gb fabric to mount files over NFS. Machines with 10Gb support have an IP address in the class C range 10.10.10.0/24 and we want to add DNS to these addresses.

1 Gb (cluster, cs)

We have two class C subnets on the 1Gb fabric: 159.28.22.0/24 (CS) and 159.28.23.0/24 (cluster). This means we have double the IP addresses on the 1Gb fabric that we have on the 10Gb fabric.

Any user accessing *.cluster.earlham.edu and *.cs.earlham.edu is making calls on a 1Gb network.

Intra-cluster fabrics

The layout cluster has an Infiniband infrastructure. Wachowski has only a 1Gb infrastructure.

Power

We have a backup power supply, with batteries last upgraded in 2019 (?). We’ve had a few outages since then and power has held up well.

HVAC

HVAC systems are static and are largely managed by Facilities.

See full topology diagrams here.

A word about what's happening between files and the drives they live on.


New sysadmins

These pages will be helpful for you if you're just starting in the group:

Note: you'll need to log in with wiki credentials to see most Sysadmin pages.

Additional information

These pages contain a lot of the most important information about our systems and how we operate.

Technical docs

Common tasks

Group and institution information