Hi, we've been using g5.8xlarge instance and could...
# ask-metaflow
f
Hi, we've been using g5.8xlarge instance and could request resources accordingly. But today we're trying and all the jobs get stuck in runnable state. We are requesting resources below the allowed compute on the instance. Any idea why it might be happening ?
1
u
Hi Harsh, to begin with, are these instances in a Kubernetes cluster? how much resources are you requesting? In K8S, there are some other services that can run on every node and thereby take up some resources. So .. a user may not be able to use all the resources on the node.
s
+1 to @User. In case you are on AWS Batch - here is a good article from AWS re: debugging workloads stuck in RUNNABLE.
f
Hi guys, no we're not using Kubernetes as of now. I was requesting gpu=1, memory=32000. We checked this article, things seem fine according the article.
a
Do you have cloud trail enabled? Happy to jump on a quick call later in the night PT if you need help debugging - lmk!
f
@square-wire-39606 Hi, we've found some sort of workaround, but would like your suggestions around our design. Can we connect in the evening PT? cc: @future-grass-38108 @acceptable-cartoon-71905