General information
Middleware
UMD
- plans on CentOS8 STARTED
CMD
- CMD-OS released http://repository.egi.eu/2020/02/18/release-cmd-os-1-3-0/
Preview repository
- released on 2020-05-08
- Preview 1.27.0 AppDB info (sl6): ARC 6.5.0 and 6.6.0, CVMFS 2.7.2, dCache 5.2.20, frontier-squid 4.11.2, gfal2 2.17.2, xrootd 4.11.3
- Preview 2.27.0 AppDB info (CentOS 7): ARC 6.5.0 and 6.6.0, CVMFS 2.7.2, dCache 5.2.20, frontier-squid 4.11.2, gfal2 2.17.2, xrootd 4.11.3
Operations
ARGO/SAM
- Swift probe included in the ARGO_MON_OPERATORS profile: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146835
- new cream probe under test: https://ggus.eu/index.php?mode=ticket_info&ticket_id=144554
- Metrics: eu.egi.CREAMCE-JobSubmit, eu.egi.CREAMCE.WN-Csh, eu.egi.CREAMCE.WN-Softver
- results: 183 endpoints, 21 WARNING (Timeout occurred (900 sec) ), 69 CRITICAL. Success rate 66.67% (50.8 % considering the WARNING)
- to include in the ARGO_MON_OPERATORS profile
- fixed a problem with eu.egi.CREAMCE.WN-Softver: now it is checked also the UMD version
When successful:
CREAM JobOutput OK: retrieved outputSandbox: ['std.err', 'std.out'] **** std.err **** + versionFilter='ή\._PIPE_ί\._PIPE_έ\._PIPE_ΰ\.' + type=unknow + mwver=error + '[' -f /etc/umd-release ']' + type=UMD ++ cat /etc/umd-release ++ awk '{print $3}' + mwver=4.1.3 + set +x **** std.out **** atlaswn184 has UMD 4.1.3
When it fails:
CREAM JobOutput ERROR [DONE-OK, exitCode=1 ]: retrieved outputSandbox: ['std.err', 'std.out'] **** std.err **** + versionFilter='ή\._PIPE_ί\._PIPE_έ\._PIPE_ΰ\.' + type=unknow + mwver=error + '[' -f /etc/umd-release ']' + '[' -f glite-version ']' + '[' -f /etc/emi-version ']' + '[' -f lcg-version ']' ++ hostname -s + echo 'ERROR: [glite_PIPE_lcg_PIPE_emi]-version was not found in n172' + exit 1 **** std.out **** ERROR: [glite_PIPE_lcg_PIPE_emi]-version was not found in n172
- HTCondor-CE probe deployed on the test instance: https://ggus.eu/index.php?mode=ticket_info&ticket_id=141177
- 96 endpoints, 3 WARNING, 14 CRITICAL, success rate is about 82.3%
- to include in the ARGO_MON_OPERATORS profile
FedCloud
Feedback from DMSU
Monthly Availability/Reliability
- Under-performed sites in the past A/R reports with issues not yet fixed:
- AsiaPacific: https://ggus.eu/index.php?mode=ticket_info&ticket_id=142591
- NGI_CH: https://ggus.eu/index.php?mode=ticket_info&ticket_id=145818
- UNIGE-DPNC: new ARC-CE put in production at the end of April, A/R figures are improving.
- NGI_IT:
- INFN-BARI: https://ggus.eu/index.php?mode=ticket_info&ticket_id=145815 SRM failures
- GARR-01-DIR: https://ggus.eu/index.php?mode=ticket_info&ticket_id=145458 failures during the DPM upgrade, debug is ongoing; the problem was with the file /etc/dmlite.conf.d/mysql.conf: it wasn't empty and the plugin wasn't activated.
- INFN-CATANIA-STACK: https://ggus.eu/index.php?mode=ticket_info&ticket_id=144715 failures with IGTF and OpenStack probes
- NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=143945
- ICM: qcg and SRM failures; statistics decreasing again;
- NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=140557
- TASK: QCG problems was fixed; some CREAM-CE failures; the SRM tests are still failing and the DPM upgrade isn't completed yet
- NGI_UK:
- UKI-NORTHGRID-SHEF-HEP: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146455 needing to re-install the ARC-CE
- UKI-SOUTHGRID-SUSX: https://ggus.eu/index.php?mode=ticket_info&ticket_id=144720 Migration from CREAM to ARC, WN migration to CentOS7; SRM to be decommissioned; ARC-CE is failing the tests (now the IGTF is failing)
- ROC_CANADA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146452
- CA-SFU-T2: SLURM problems caused failures to site-BDII freshness check due to some old jobs not properly cancelled; A/R figures are now improving
- CA-WATERLOO-T2: SRM failures not involving production VOs, fixed; some unscheduled downtime affected the the A/R figures
- ROC_LA https://ggus.eu/index.php?mode=ticket_info&ticket_id=145812
- GRID-UNAM: DPM upgrade not yet completed;
- NGI_UA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146454
- UA-KNU: GlueSEUniqueID wasn't published due to missing DNS back-resolving record for IPv6 address of the SE. Failures on ARC-CEs then fixed.
- Under-performed sites after 3 consecutive months, under-performed NGIs, QoS violations: (April 2020):
- AfricaArabia: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146877
- ZA-WITS-CORE: SE hardware problem, machine sent to the vendor; CREAM-CE failures
- NGI_AEGIS: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146867
- AEGIS03-ELEF-LEDA: SRM problems, endpoint put out of production until the issues are fixed
- NGI_CH: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146870
- CSCS-LCG2
- UNIBE-LHEP
- NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146871
- GoeGRID: CREAM-CE issues have been solved, failures with ARC-CE
- NGI_IT: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146868
- CIRMMP
- INFN-LECCE
- NGI_NDGF: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146876
- UNICPH-NBI
- NGI_NL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146872 (SOLVED)
- SARA-MATRIX: CAs packages not properly updated, fixed
- NGI_UA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146874
- UA-IFBG: ARC-CE problems, fixed
- UA-IRE: ARC-CE issues, fixed
- UA_BITP_ARC: ARC-CE mis-configuration, fixed
- NGI_UK: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146875
- UKI-LT2-QMUL
- ROC_LA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146873
- SUPERCOMPUTO-UNAM
- AfricaArabia: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146877
- sites suspended:
- RO-02-NIPNE: it didn't recovered from the long downtime
IPv6 readiness plans
- please provide updates to the IPv6 assessment (ongoing) https://wiki.egi.eu/w/index.php?title=IPV6_Assessment
- if any relevant, information will be summarised at OMB
ARC Middleware 5 end of support, migration to ARC 6
- No new feature development is planned or going on for ARC5 and no bug-fixing development will happen on ARC5 code base in the future except for security issues.
- Security fixes for ARC5 will be provided till end of June 2020.
- Production Sites already running ARC 5 will be able to get deployment and configuration troubleshooting help via GGUS till end June 2021. This we call "operational site support".
- ARC5 is available in EPEL7 and will stay there. EPEL8 will only contain ARC 6.
- PROC16 Decommission of unsupported software
- Number of ARC-CE endpoints registered on GOCDB (monitored and production): 138
- From the BDII:
$ ldapsearch -x -LLL -H ldap://egee-bdii.cnaf.infn.it:2170 -b "GLUE2GroupID=grid,o=glue" '(&(objectClass=GLUE2Endpoint)(&(GLUE2EndpointImplementationName=nordugrid-arc)(GLUE2EndpointTechnology=gridftp)))' GLUE2EndpointImplementationVersion GLUE2EndpointID | grep GLUE2EndpointImplementationVersion | sort | uniq -c 1 GLUE2EndpointImplementationVersion: 20190403020701 1 GLUE2EndpointImplementationVersion: 20200217020714 1 GLUE2EndpointImplementationVersion: 20200226020744 1 GLUE2EndpointImplementationVersion: 20200228020731 1 GLUE2EndpointImplementationVersion: 20200424020718 1 GLUE2EndpointImplementationVersion: 20200505020715 1 GLUE2EndpointImplementationVersion: 4.1.0 1 GLUE2EndpointImplementationVersion: 5.0.2 1 GLUE2EndpointImplementationVersion: 5.1.3 1 GLUE2EndpointImplementationVersion: 5.3.0 5 GLUE2EndpointImplementationVersion: 5.3.1 6 GLUE2EndpointImplementationVersion: 5.4.1 11 GLUE2EndpointImplementationVersion: 5.4.2 7 GLUE2EndpointImplementationVersion: 5.4.3 47 GLUE2EndpointImplementationVersion: 5.4.4 2 GLUE2EndpointImplementationVersion: 6.2.0 5 GLUE2EndpointImplementationVersion: 6.4.1 14 GLUE2EndpointImplementationVersion: 6.5.0 5 GLUE2EndpointImplementationVersion: 6.6.0
LCGDM end of support and migration to / enabling of DOME
- The DPM team has agreed to extend support for security updates for the DPM legacy functionality until 30 September 2019. However, affected service providers should still plan to disable legacy mode well before this date.
- more details: https://wiki.egi.eu/wiki/DPM_End_of_legacy-mode_support
- since DPM 1.10.3 release, it is possible enabling the non-legacy mode DOME (Disk operations Management Engine, see documentation)
- latest DPM/dmlite release 1.13.2 has been also released in UMD 4.10.0
- EGI Operations sent a broadcast to site-admins(first bunch and second bunch) and VO managers with a survey to fill in:
- site-admins: to know the upgrade plans: https://www.surveymonkey.com/r/2PBZSNB
- VO Managers: to know any barriers in moving away from using SRM protocol https://www.surveymonkey.com/r/2ZMZJVG
- Deployment statistics (May 8th):
$ ldapsearch -x -LLL -H ldap://egee-bdii.cnaf.infn.it:2170 -b "GLUE2GroupID=grid,o=glue" '(&(objectClass=GLUE2Manager)(GLUE2ManagerProductName=DPM))' GLUE2ManagerProductVersion GLUE2ManagerID | grep GLUE2ManagerProductVersion | sort | uniq -c 1 GLUE2ManagerProductVersion: 1.10.0 66 GLUE2ManagerProductVersion: 1.13.0 2 GLUE2ManagerProductVersion: 1.13.1 12 GLUE2ManagerProductVersion: 1.13.2 3 GLUE2ManagerProductVersion: 1.8.10 1 GLUE2ManagerProductVersion: 1.8.8 1 GLUE2ManagerProductVersion: 1.8.9 4 GLUE2ManagerProductVersion: 1.9.0
Liasing with WLCG to follow-up the upgrade. Opened GGUS tickets asking the following:
- all the sites with older DPM versions than 1.12 are suggested to upgrade to the latest DPM version , following the guide DPM upgrade (chapter 1 Upgrade to DPM 1.10.0 "Legacy Flavour" and chapter 2 Upgrade to DPM 1.10.0 "Dome Flavour")
- DOME and the old LCGDM (srm protocol) will coexist
- Monitoring: sites should enable the monitoring of the HTTP/WebDav and/or GridFTP endpoints
- register the storage service endpoint as webdav and/or globus-GRIDFTP service type, with production flag disabled, providing respectively the URL field and the Extension Properties information as explained in the HOWTO21
- check if the tests are ok
- switch the production flag to "yes"
List of tickets
- DPM upgrade, DOME enabling, and monitoring (33)
- enabling DOME and proper monitoring (6)
- tickets opened by WLCG: 13 and 7
SECMON failures
Several CEs are failing the job submission tests, preventing pakiti to check the vulnerabilities fixes on the WNs.
- original ticket: https://ggus.eu/index.php?mode=ticket_info&ticket_id=143837
- List of tickets to the sites
- https://ggus.eu/index.php?mode=ticket_info&ticket_id=144732
AOB
- Storage accounting: http://goc-accounting.grid-support.ac.uk/storagetest/storagesitesystems.html
- several sites stopped publishing storage accounting records: it needs to investigate on and fix it
Next meeting
June 8th, 2020 https://indico.egi.eu/event/4900/