#12620 Post datacenter move work
Closed: Fixed with Explanation by kevin. Opened by kevin.

This ticket is collecting the items from the datacenter move hackmd document and putting them in a better place for people to work on.

Larger items can (and will) be split off on their own tickets if need be.

  • power10 reconfiguration - This needs investigation, we need to decide how well a vHMC will help us and if it's worth it. @kevin and @arrfab are investigating
  • Move to sigul for secure-boot signing - we have much of this all ready, just need to work on setting up siguldry on a bvmhost to proxy things.
  • repurpose the host I had set for autosign02 to be buildhw-x86-04
  • Still some openshift apps to deploy in staging (see hackmd for full list)
  • Still some openshift apps to deploy in prod (see hackmd)
  • mailman01 is OOM killing mailman3 process (I set it to autorestart on fail now)
  • Some playbooks are not completing, in particular the releng-compose playbook.
  • We still need to re-enable backups next week after things have stablized
  • We still need to carefully shut down iad2 hardware (setting single user mode default, powering off, etc)

There may be more small things i missed.


  • flatpak builds don't work because the login-registry role is being skipped on buildvm's (at least that seems likely the cause)

Just adding hackmd document link https://hackmd.io/54xmtW6IQoKNKnbRXySxSg

Metadata Update from @zlopez:
- Issue priority set to: Waiting on Assignee (was: Needs Review)
- Issue tagged with: high-gain, high-trouble, ops

One more thing I just thought of: some mirrors may have only allowed our previous ip/range to do rsync for checking the mirror freshness. We may have to check that manually in the logs and mirrormanager db.
I've sent an email to the mirrors-admin list, but it's not a guarantee that mirror owners will update their ACLs.
A quick grep in the logs show 18 mirrors with rsync connection refused errors.

Shouldn't there be a cron job error sent as well by e-mail?

Another one to add

  • zabbix playbook is failing

I assume the ssh configuration in https://docs.fedoraproject.org/en-US/infra/sysadmin_guide/sshaccess/ should be updated?

We need to re-enroll all our non rdu3 hosts... this I think needs a ansible run on them (to change /etc/hosts entries) and then the uninstall/install ipa client dance and fixing sssd.conf, etc.
I already fixed people01 I think.

I assume the ssh configuration in https://docs.fedoraproject.org/en-US/infra/sysadmin_guide/sshaccess/ should be updated?

Yeah, that change was merged, but docsbuilding wasn't working. Should work later today hopefully.

Fedocal login isn't working for me, it's a mix of 400 errors or bouncing back to the main page.

Calendar login page tries to send you to:

https://calendar.fedoraproject.org/login/?next=https://calendar.fedoraproject.org,calendar.fedoraproject.org/

...and then id.fp.o gets upset and gives "400 Bad request: redirect_uri" fwiw you can make the id.fp.o work by manually changing URLs, but then calendar thinks weird things are going on and refuses to accept the auth.

pkgs01.rdu3: /var/log/messages and secure are empty again ... something to do with log rollover isn't happy.

Everything is fine after I restart rsyslog, but all the logs between the rollover and the restart are gone.

Iad2 hardware shutdown can be tracked in https://pagure.io/fedora-infrastructure/issue/12625 now.

Retrospective hackmd collection: https://hackmd.io/wB6rboFvS3uCErEOpuzp4g

Copr login now goes into an infinite redirect loop. Not sure if related.

ELN composes are happening but aren't getting pushed to mirrors; https://dl.fedoraproject.org/pub/eln/1/ hasn't been updated since the move.

Copr login now goes into an infinite redirect loop. Not sure if related.

See #12623 (TLDR: use the oidc login button, not the login button)

ELN composes are happening but aren't getting pushed to mirrors; https://dl.fedoraproject.org/pub/eln/1/ hasn't been updated since the move.

Fixed. Was missing the /pub mount. The next compose should sync.

Some broken cron jobs:

  • get_retired_packages.sh in releng repo gets a 'parse error: Invalid numeric literal at line 1, column 10' CC: @lenkaseg

  • ipa02.stg/ipa03.stg fail backups with:

/etc/cron.daily/data-only-backup.sh:

ls: cannot access '/var/lib/ipa/backup/ipa-data-*': No such file or directory
Error: Local roles CA do not match globally used roles CA, KRA. A backup done on this host would not be complete enough to restore a fully functional, identical cluster.
The ipa-backup command failed. See /var/log/ipabackup.log for more information

CC: @zlopez 
- condense-mirrorlogs cron is failing with a long error... CC: @james  or perhaps @nphilipp ?
- package-owner-alias cron is failing on bastion01/02... ends in: 

simplejson.errors.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
error creating owner-alias file
```

The page https://src.fedoraproject.org/ssh_info needs updating too, it still lists the old ssh keys.

Calendar login page tries to send you to:

https://calendar.fedoraproject.org/login/?next=https://calendar.fedoraproject.org,calendar.fedoraproject.org/

...and then id.fp.o gets upset and gives "400 Bad request: redirect_uri" fwiw you can make the id.fp.o work by manually changing URLs, but then calendar thinks weird things are going on and refuses to accept the auth.

This one is fixed, it required a couple things:
- code change to better handle the proxy headers
- base image change (it was set to run on python 3.6)
- code change to adapt to changes in newer versions of Flask and other deps
- OIDC config change to adapt to the new callback URL in recent version of flask-oidc

Metadata Update from @zlopez:
- Issue tagged with: dc-move

Cannot add ipa groups:

IPA Error 4203: DatabaseError
Operations error: Allocation of a new value for range cn=posix ids,cn=distributed numeric assignment plugin,cn=plugins,cn=config failed! Unable to proceed.

ELN composes sync fixed.
meetbot-raw is back and fixed.

Also, new users cannot register:

ERROR in registration: An unhandled error BadRequest happened while activating stage user REDACTED: Operations error: Allocation of a new value for range cn=posix ids,cn=distributed numeric assignment plugin,cn=plugins,cn=config failed! Unable to proceed.

Also, new users cannot register:

ERROR in registration: An unhandled error BadRequest happened while activating stage user REDACTED: Operations error: Allocation of a new value for range cn=posix ids,cn=distributed numeric assignment plugin,cn=plugins,cn=config failed! Unable to proceed.

Fixed: https://pagure.io/fedora-infrastructure/issue/12641

I fixed the cron.job for staging IPA deployment.

Not sure what happened, but the iad2 servers were still there in ipa02.stg and ipa03.stg. So I just removed them from the topology.

I will try to look at the production IPA groups and users next.

EDIT: I see that the issue is already fixed. Let me check the next one on the list :-)

Hi team! Is http://data-analysis.fedoraproject.org/ down because of this? or has it just moved address? I was trying to look up some countme data. thank you very much! Found this through https://status.fedoraproject.org/

So a bunch of things fixed:

[x] - Some playbooks are not completing, in particular the releng-compose playbook.
[x] - We still need to carefully shut down iad2 hardware (setting single user mode default, powering off, etc)
[x] - ssh configuration in https://docs.fedoraproject.org/en-US/infra/sysadmin_guide/sshaccess/ updated
[x] - All non rdu3 hosts should be re-enrolled in rdu3 ipa
[x] - fedocal login should work
[x] - ELN composes are syncing out
[x] - data-analysis was pointing to the old DC, I updated it. Should start working as soon as cache times out.

Oustanding issues:

Will move to their own tickets:

power10 reconfig
sigul secure-boot signing
reconfigure autosign02

Outstanding:

flatpak builds not working. Is this still happening?

zabbix playbook is failing

get_retired_packages.sh in releng repo gets a 'parse error: Invalid numeric literal at line 1, column 10'

condense-mirrorlogs cron is failing with a long error...

package-owner-alias cron is failing on bastion01/02

thank you very much @kevin !

Another random thing: sundries01.stg.rdu3.fedoraproject.org is blocking the nightly ansible playbook checks by putting the task into D state.

I think thats due to a lingering mtu 9000 on the vm... a 'nmcli c up eth0' should reset it back to 1500.

Sorry off-topic but FYI anyone who's looking at the countme data, this week the the server migration lead to no weekly-active-users being counted for a few days, and so the weekly-active-users will be around 40% lower for this week across the board. partial data week!

Let's not forget that there is a lot of alerts in nagios as well https://nagios.fedoraproject.org/nagios/

Hi @kevin , do you still see the get_retired_packages.sh in releng repo failing?
I see the output correctly displayed on the lookaside and @zlopez run the script and it seemed not to fail. Maybe it was just some DC move related hiccup?

Here's the last failed one I see:

https://lists.fedoraproject.org/archives/list/releng-cron@lists.fedoraproject.org/message/Z2FDRVW34KYLTDPKELEAGBMOUFWZDI2D/

get_retired_packages.sh: line 18: pushd: /srv/git/rpms/*.git: No such file or directory

should that be perhaps /srv/git/repositories/rpms/*.git ?
But it shouldn't have changed. ;(

Another thing I found today: vmhost-p09-copr01.rdu-cc.fedoraproject.org has set this URL as repository for fedora https://infrastructure.fedoraproject.org/pub/archive/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml, which returns 404.

I checked batcave01 and archive folder is really empty, but accessing it without archive folder works like this https://infrastructure.fedoraproject.org/pub/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml. I wonder if the archive folder should be some kind of symlink or the URL for repo is wrong.

I can’t find it here yet, but we talked about it in the team before: On pkgs01, the httpd process regularly dumps core, probably when a worker process winds down to be restarted. Others and I have looked into it, ~~I’ll create a ticket with the info I know and link here once it’s there.~~ I’ve created #12670 with some info I collected over the day.

Another thing I found today: vmhost-p09-copr01.rdu-cc.fedoraproject.org has set this URL as repository for fedora https://infrastructure.fedoraproject.org/pub/archive/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml, which returns 404.

I checked batcave01 and archive folder is really empty, but accessing it without archive folder works like this https://infrastructure.fedoraproject.org/pub/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml. I wonder if the archive folder should be some kind of symlink or the URL for repo is wrong.

Yeah, we should just make sure and move this machine to a supported release. I mean, we can add archive, but... we shouldn't.

power10 reconfig: https://pagure.io/fedora-infrastructure/issue/12674
secure boot signing changes can be tracked in: https://pagure.io/fedora-infrastructure/issue/7361
autosign02 reconfigure: I just did this.
flatpak builds not working: https://pagure.io/fedora-infrastructure/issue/12675
package-owner-alias cron seems working now.
get_retired_packages: https://pagure.io/fedora-infrastructure/issue/12676
package_owner_alias seems working/fixed.

I'm going to close this ticket now, if there's anything I missed, please do file it as a serpate ticket. Thanks everyone!

Metadata Update from @kevin:
- Issue close_status updated to: Fixed with Explanation
- Issue status updated to: Closed (was: Open)

Yeah, we should just make sure and move this machine to a supported release. I mean, we can add archive, but... we shouldn't.

Finished, I updated it to F42 and did some tweaks to ansible to get the playbook running successfully.

Metadata