updating to latest Kubernetes version 1.20.0 fix
security issues. In the current version, it allows
for the token leak in logs when logLevel >= 9.
Signed-off-by: subhamkrai <srai@redhat.com>
Currently, mgr deployment has two overlapping environment variables, ROOK_POD_IP.
This causes kube-apiserver to create error logs
Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
The arbiter can only be configured with the stretch cluster if the
CRUSH map is balanced and there are two zones in the CRUSH map.
After the OSDs are configured, we wait for all the OSD pods to be
running and that the CRUSH map is balanced. If it takes more than
two minutes, we fail the reconcile and try again. This is only done
the first time the stretch cluster is configured. In future reconciles
we first check if the stretch cluster is already enabled before
enabling it again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This adds the functionality to add custom pod labels to the CSI
components through the operator configuration way of env vars or config
map.
Resolves#6593
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
If users want to restrict the nodes where Ceph daemons should exist,
it's better to make the placement of admission controller configurable
as other daemons.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
Instead of directly defining our own constants, we should be using the k8s
constants for the well-known topology labels.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
this commit handle golangci-lint linter errcheck.
`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases
To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit handle golangci-lint linter staticcheck error.
`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.
To see only `staticcheck` linter output
`golangci-lint run --disable-all -E staticcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
If the pod spec changed, we expect an upgrade to proceed for that daemon.
If the check for a changed pod spec fails, we were skipping the update
of that daemon. Instead of skipping the update, we now assume the pod
spec changed if we fail to detect the change so that we can ensure the upgrade
even if we check for upgrades too often.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In the whereabouts config, the network range is determined by range
option unlike other cni plugins that use Subnet option.
Signed-off-by: rohan47 <rohgupta@redhat.com>
The name generation utils need to provide consistent outputs across
versions so that names generated based on values are consistent and
older generated keys are not lost.
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
All assignments made to errors or other variables should be
used instead of ignored. This checks for errors in a couple
places that were previously ignored, and also initializes
a variable such that it won't be always overwritten by the
various if else statements.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add the ability to provision Ceph OSDs with Drive Groups.
This adds Drive Groups to the CephCluster CRD, and it sets code
in place for propagating this config to the OSD provisioning pod.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
remove csi drivers and delete k8s
services that drivers create,when
the default setting of drivers changed
to false.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
The operator should only print helpful info messages
when an OSD is going to be updated. If the OSD hasn't
changed it is just a debug message.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The pod deletion is needed by other daemons besides the osds,
so we move the helper method into the k8sutil package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
An OSD on a PVC when portable=false is assigned to a node
with a node selector. The same node assignment is expected
for the lifetime of the cluster. On subsequent reconciles,
the operator was looking up the node assignment from the
nodeName on the pod spec, which is not set if the pod is down.
The operator needs to retrieve the assignment from the
deployment spec node selector.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator settings can come either from the configmap, an
env var, or a default value. Knowing where the setting is loaded
from is an important detail to see in the log.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
YamlToContainerResource can be used to convert the raw
string data to the array of ContainerResource.
the ContainerResource can be applied to the
pod container resources request and limits.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:
network:
provider: multus (or 'host' for hostNetworking)
selectors:
public: NetworkAttachmentDefinition name
cluster: NetworkAttachmentDefinition name
If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.
Please refer to the doc to learn more.
Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.
Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.
Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>
with older implementation, servicemonitor was not getting
updated due to missing resource version. This fix adds
resource version to the servicemonitor definition and
ensures that the object is properly created or
updated.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.
Signed-off-by: Sébastien Han <seb@redhat.com>