This ticket is collecting the items from the datacenter move hackmd document and putting them in a better place for people to work on.
Larger items can (and will) be split off on their own tickets if need be.
There may be more small things i missed.
Just adding hackmd document link https://hackmd.io/54xmtW6IQoKNKnbRXySxSg
Metadata Update from @zlopez: - Issue priority set to: Waiting on Assignee (was: Needs Review) - Issue tagged with: high-gain, high-trouble, ops
One more thing I just thought of: some mirrors may have only allowed our previous ip/range to do rsync for checking the mirror freshness. We may have to check that manually in the logs and mirrormanager db. I've sent an email to the mirrors-admin list, but it's not a guarantee that mirror owners will update their ACLs. A quick grep in the logs show 18 mirrors with rsync connection refused errors.
Shouldn't there be a cron job error sent as well by e-mail?
Another one to add
I assume the ssh configuration in https://docs.fedoraproject.org/en-US/infra/sysadmin_guide/sshaccess/ should be updated?
We need to re-enroll all our non rdu3 hosts... this I think needs a ansible run on them (to change /etc/hosts entries) and then the uninstall/install ipa client dance and fixing sssd.conf, etc. I already fixed people01 I think.
Yeah, that change was merged, but docsbuilding wasn't working. Should work later today hopefully.
Fedocal login isn't working for me, it's a mix of 400 errors or bouncing back to the main page.
Calendar login page tries to send you to:
https://calendar.fedoraproject.org/login/?next=https://calendar.fedoraproject.org,calendar.fedoraproject.org/
...and then id.fp.o gets upset and gives "400 Bad request: redirect_uri" fwiw you can make the id.fp.o work by manually changing URLs, but then calendar thinks weird things are going on and refuses to accept the auth.
pkgs01.rdu3: /var/log/messages and secure are empty again ... something to do with log rollover isn't happy.
Everything is fine after I restart rsyslog, but all the logs between the rollover and the restart are gone.
restart rsyslog
Iad2 hardware shutdown can be tracked in https://pagure.io/fedora-infrastructure/issue/12625 now.
Retrospective hackmd collection: https://hackmd.io/wB6rboFvS3uCErEOpuzp4g
Copr login now goes into an infinite redirect loop. Not sure if related.
ELN composes are happening but aren't getting pushed to mirrors; https://dl.fedoraproject.org/pub/eln/1/ hasn't been updated since the move.
See #12623 (TLDR: use the oidc login button, not the login button)
Fixed. Was missing the /pub mount. The next compose should sync.
Some broken cron jobs:
get_retired_packages.sh in releng repo gets a 'parse error: Invalid numeric literal at line 1, column 10' CC: @lenkaseg
ipa02.stg/ipa03.stg fail backups with:
/etc/cron.daily/data-only-backup.sh:
ls: cannot access '/var/lib/ipa/backup/ipa-data-*': No such file or directory Error: Local roles CA do not match globally used roles CA, KRA. A backup done on this host would not be complete enough to restore a fully functional, identical cluster. The ipa-backup command failed. See /var/log/ipabackup.log for more information
CC: @zlopez - condense-mirrorlogs cron is failing with a long error... CC: @james or perhaps @nphilipp ? - package-owner-alias cron is failing on bastion01/02... ends in:
simplejson.errors.JSONDecodeError: Expecting value: line 1 column 1 (char 0) error creating owner-alias file ```
The page https://src.fedoraproject.org/ssh_info needs updating too, it still lists the old ssh keys.
Calendar login page tries to send you to: https://calendar.fedoraproject.org/login/?next=https://calendar.fedoraproject.org,calendar.fedoraproject.org/ ...and then id.fp.o gets upset and gives "400 Bad request: redirect_uri" fwiw you can make the id.fp.o work by manually changing URLs, but then calendar thinks weird things are going on and refuses to accept the auth.
This one is fixed, it required a couple things: - code change to better handle the proxy headers - base image change (it was set to run on python 3.6) - code change to adapt to changes in newer versions of Flask and other deps - OIDC config change to adapt to the new callback URL in recent version of flask-oidc
Metadata Update from @zlopez: - Issue tagged with: dc-move
Cannot add ipa groups:
IPA Error 4203: DatabaseError Operations error: Allocation of a new value for range cn=posix ids,cn=distributed numeric assignment plugin,cn=plugins,cn=config failed! Unable to proceed.
ELN composes sync fixed. meetbot-raw is back and fixed.
Also, new users cannot register:
ERROR in registration: An unhandled error BadRequest happened while activating stage user REDACTED: Operations error: Allocation of a new value for range cn=posix ids,cn=distributed numeric assignment plugin,cn=plugins,cn=config failed! Unable to proceed.
Also, new users cannot register: ERROR in registration: An unhandled error BadRequest happened while activating stage user REDACTED: Operations error: Allocation of a new value for range cn=posix ids,cn=distributed numeric assignment plugin,cn=plugins,cn=config failed! Unable to proceed.
Fixed: https://pagure.io/fedora-infrastructure/issue/12641
I fixed the cron.job for staging IPA deployment.
Not sure what happened, but the iad2 servers were still there in ipa02.stg and ipa03.stg. So I just removed them from the topology.
I will try to look at the production IPA groups and users next.
EDIT: I see that the issue is already fixed. Let me check the next one on the list :-)
Hi team! Is http://data-analysis.fedoraproject.org/ down because of this? or has it just moved address? I was trying to look up some countme data. thank you very much! Found this through https://status.fedoraproject.org/
So a bunch of things fixed:
[x] - Some playbooks are not completing, in particular the releng-compose playbook. [x] - We still need to carefully shut down iad2 hardware (setting single user mode default, powering off, etc) [x] - ssh configuration in https://docs.fedoraproject.org/en-US/infra/sysadmin_guide/sshaccess/ updated [x] - All non rdu3 hosts should be re-enrolled in rdu3 ipa [x] - fedocal login should work [x] - ELN composes are syncing out [x] - data-analysis was pointing to the old DC, I updated it. Should start working as soon as cache times out.
Oustanding issues:
Will move to their own tickets:
power10 reconfig sigul secure-boot signing reconfigure autosign02
Outstanding:
flatpak builds not working. Is this still happening?
zabbix playbook is failing
get_retired_packages.sh in releng repo gets a 'parse error: Invalid numeric literal at line 1, column 10'
condense-mirrorlogs cron is failing with a long error...
package-owner-alias cron is failing on bastion01/02
thank you very much @kevin !
Another random thing: sundries01.stg.rdu3.fedoraproject.org is blocking the nightly ansible playbook checks by putting the task into D state.
sundries01.stg.rdu3.fedoraproject.org
I think thats due to a lingering mtu 9000 on the vm... a 'nmcli c up eth0' should reset it back to 1500.
Sorry off-topic but FYI anyone who's looking at the countme data, this week the the server migration lead to no weekly-active-users being counted for a few days, and so the weekly-active-users will be around 40% lower for this week across the board. partial data week!
Let's not forget that there is a lot of alerts in nagios as well https://nagios.fedoraproject.org/nagios/
Hi @kevin , do you still see the get_retired_packages.sh in releng repo failing? I see the output correctly displayed on the lookaside and @zlopez run the script and it seemed not to fail. Maybe it was just some DC move related hiccup?
get_retired_packages.sh
Here's the last failed one I see:
https://lists.fedoraproject.org/archives/list/releng-cron@lists.fedoraproject.org/message/Z2FDRVW34KYLTDPKELEAGBMOUFWZDI2D/
get_retired_packages.sh: line 18: pushd: /srv/git/rpms/*.git: No such file or directory
should that be perhaps /srv/git/repositories/rpms/*.git ? But it shouldn't have changed. ;(
Another thing I found today: vmhost-p09-copr01.rdu-cc.fedoraproject.org has set this URL as repository for fedora https://infrastructure.fedoraproject.org/pub/archive/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml, which returns 404.
vmhost-p09-copr01.rdu-cc.fedoraproject.org
fedora
https://infrastructure.fedoraproject.org/pub/archive/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml
I checked batcave01 and archive folder is really empty, but accessing it without archive folder works like this https://infrastructure.fedoraproject.org/pub/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml. I wonder if the archive folder should be some kind of symlink or the URL for repo is wrong.
batcave01
archive
https://infrastructure.fedoraproject.org/pub/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml
I can’t find it here yet, but we talked about it in the team before: On pkgs01, the httpd process regularly dumps core, probably when a worker process winds down to be restarted. Others and I have looked into it, ~~I’ll create a ticket with the info I know and link here once it’s there.~~ I’ve created #12670 with some info I collected over the day.
pkgs01
httpd
Another thing I found today: vmhost-p09-copr01.rdu-cc.fedoraproject.org has set this URL as repository for fedora https://infrastructure.fedoraproject.org/pub/archive/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml, which returns 404. I checked batcave01 and archive folder is really empty, but accessing it without archive folder works like this https://infrastructure.fedoraproject.org/pub/fedora-secondary/updates/40/Everything/ppc64le/repodata/repomd.xml. I wonder if the archive folder should be some kind of symlink or the URL for repo is wrong.
Yeah, we should just make sure and move this machine to a supported release. I mean, we can add archive, but... we shouldn't.
power10 reconfig: https://pagure.io/fedora-infrastructure/issue/12674 secure boot signing changes can be tracked in: https://pagure.io/fedora-infrastructure/issue/7361 autosign02 reconfigure: I just did this. flatpak builds not working: https://pagure.io/fedora-infrastructure/issue/12675 package-owner-alias cron seems working now. get_retired_packages: https://pagure.io/fedora-infrastructure/issue/12676 package_owner_alias seems working/fixed.
I'm going to close this ticket now, if there's anything I missed, please do file it as a serpate ticket. Thanks everyone!
Metadata Update from @kevin: - Issue close_status updated to: Fixed with Explanation - Issue status updated to: Closed (was: Open)
Finished, I updated it to F42 and did some tweaks to ansible to get the playbook running successfully.