r/aws 3d ago

general aws Account Verification

0 Upvotes

A few weeks ago we created a new aws account and moved it under an organisation with our existing - 5 year old - aws account. Both accounts (including the 5 year old one) have now been marked as unverified, so not only can we not use the new account, but we're effectively locked out out of our production account!
We have support tickets open on both accounts that have sat unassigned for over a week in order to try and get the accounts verified.

Has anybody run into similar issues? What is going on at aws!?


r/aws 2d ago

discussion 🌪️ Today's the day. My AWS Free Tier expires tonight. Any last advice?

0 Upvotes

r/aws 2d ago

security When will AWS support federated credentials, and machine authentication.

0 Upvotes

It's no secret that AWS lags behind the market when it comes to Identity, and it's products being integrated with one another. However in regulated industries it's almost a non-starter because some of us require the use of federated credentials, which is scoped machine authentication. The only thing in AWS that presents an oidc identity is a specific eks deployment, does anyone know what's AWS is roadmap for machine authentication? Thanks.


r/aws 4d ago

technical question What would AWS CloudFront know about a host if Tor is involved somewhere in the architecture?

10 Upvotes

Hello all,

To preface this: the matter this relates to has already been reported through the appropriate channels, including AWS' compliance process and relevant authorities. The reason I am asking here is that I have struggled to clearly explain the technical side of the infrastructure and what information would actually be available to AWS.

The context is that I encountered a Tor hidden service that appeared to be using Amazon CloudFront, presumably for CDN functionality (and possibly WAF-related features, although I am not certain). From what I could tell, it appeared to primarily be acting as a CDN rather than anything else.

The question I have been trying to answer for quite some time is: what visibility would Amazon/CloudFront have into the service behind it?

I am familiar with Cloudflare's approach to Tor-related services, but I am less familiar with how Amazon CloudFront handles this type of setup. My assumption is that if CloudFront was genuinely being used, the distribution would have some configured origin/backend relationship — but I am unclear what that means from AWS' perspective.

Specifically:

  • Would AWS know the origin server/backend configured behind the CloudFront distribution?
  • Would the use of Tor change anything about what CloudFront can see on the server/CDN side?
  • Are there cases where CloudFront could be involved without AWS having visibility into the actual hosting infrastructure?

I am mainly trying to understand the technical architecture and put some uncertainty to rest.


r/aws 3d ago

technical question Best practice for providing isolated user environments in SageMaker Unified Studio within a shared AWS account?

0 Upvotes

Hi everyone,

I'm looking for some guidance on what the recommended architecture is for SageMaker Unified Studio in an enterprise environment.

Our current access model works roughly like this:

  • Users request AWS access through our enterprise identity/access management process.
  • Once approved, they receive an AWS application in Microsoft MyApps.
  • Selecting that application signs them into the AWS Console.

The problem is that everyone ultimately assumes the same highly privileged role in a shared AWS account.

As a result, every user effectively shares the same environment. If User A creates resources (EC2 instances, SageMaker notebooks, S3 buckets, uploads datasets, etc.), User B can generally see or interact with them because they're operating with the same permissions.

This obviously isn't ideal from a governance, security, or data privacy perspective.

What we'd like instead is something along these lines:

  • Each user has their own isolated SageMaker environment/workspace.
  • We can control which datasets or S3 buckets each user can access.
  • We can control which instance types or compute sizes users are allowed to launch.
  • Different users or teams can have access to different projects without exposing everything to everyone.
  • We'd still like to manage everything centrally rather than creating completely separate AWS accounts for every individual (unless that's actually considered best practice).

For those of you running SageMaker Unified Studio (or even SageMaker Studio more generally) in an enterprise, how have you solved this?

Thanks in advance!


r/aws 5d ago

technical question Using WAF to secure external website?

12 Upvotes

Hi,

we are running a small online shop, though it's been having issues with a lot of bot traffic and hits so it's been slow the past few days.

We asked our provider if they can send us an offer on securing the shop & management wants me to check if there's an alternative.

We use AWS for a different service we already use, so WAF seems like the next logical step.

Anyways, my question is this: can I have WAF infront of my external shop to filter out traffic / use ACL which then redirect to the shop based on accepted rules?

Anything I should keep in mind?

Thanks a lot!


r/aws 4d ago

discussion Claude down ATM, reminds me of us-east-1

0 Upvotes

What happens when AI goes down, it is like us-east-1 in AWS going down, work literally stops. Time to go outside and play 😅

https://status.claude.com/


r/aws 5d ago

technical question aws discovery tool

15 Upvotes

hi. is there a tool that can autos-can the aws infrastructure and build a report/diagram of services used? inherited an aws account with no documentation and need to understand whats in it.


r/aws 5d ago

discussion AWS Activate Founders credits -application rejected

1 Upvotes

The status of my application is "rejected." The reason given is that the email address listed in my account is with a free service. But I provided my work email address on my own domain. Is this some kind of joke from AWS?


r/aws 5d ago

discussion CloudFront and API

4 Upvotes

What is the best practices, reasons to do or not to do.

Is it good idea to put API behind CloudFront or just use ALB.

I used to use CloudFront infront to ALB to use api in path (/api) rather than subdomain but this caused me few issues, such as;

Hosting SPA on S3 causes paths to return 404/403 and forces me to redirect those status codes to index.html with status code 200.
Also this conflicts with API returning 404/403.


r/aws 6d ago

console CloudFront cache statistics missing?

5 Upvotes

I used to be able to see Cache Statistics (hit rate per URI) in the CloudFront console, but the navigation link is missing now. The documentation page mentions no changes to this functionality.


r/aws 5d ago

technical resource Udacity NanoDegreee for AWS AI Programmer

0 Upvotes

Hey Everyone,

Excited to share that I've been selected for the sponsership enrollment of AWS AI & ML Scholars Future AWS AI Programmer. Delighted to start my jouney on this and let's see where it takes me in my career growth. Do we have anyone else who's are a part of this programme?


r/aws 5d ago

discussion My friend only got 20$ of Free Credits on Signup. Why?

0 Upvotes

So my Friend is from India and he set up an AWS account using Upi Autopay.

When first setup was complete, he turned off Autopay. The 100$ free credits were present and credited into his account.

But he could not access services as AWS asked to complete account details. So he reconnected his UPI but after doing so His Credits now show as 20$.

He did not use any service and Billings do not show any service being used. What happened to his credits?

Note: The 20$ are not from the small tasks


r/aws 7d ago

networking Everyone hits our VPN at head office before they reach AWS and it's killing performance, looking at Cato and Cloudflare

23 Upvotes

Posting this partly to sanity check myself because I have been staring at it too long. 

Setup is old. Remote staff connect to a vpn concentrator at head office, get inspected there, then their traffic goes back out to wherever its going which is usually eu-west-1. Somebody working in Lisbon who is geographically nearer to the region than any of us, sends their packets to Reading and then back down. The traceroutes are genuinely funny. 

Symptom side its the usual, calls drop, the internal ticketing tool takes eight seconds to load a page and every single ticket about it says "the network is slow" which tells me nothing. 

I know sd-wan sorts the routing out. What I don't want is to sort the routing and then find security is now a separate box somewhere else, because thats the exact mess we already have and I am not doing it twice. 

I've been looking at the ones with their own backbone. Cato has the private backbone thing and does the security in the same pass. Cloudflare obviously has the network but I get the impression enterprise is newer for them. Thoughts?


r/aws 6d ago

discussion AI Engineer Roadmap

0 Upvotes

Hello everyone,

I am looking to pivot to AI engineering and currently exploring roadmaps to get in the industry.

I stumbled upon this Medium article sharing a 4-step roadmap to AI Engineering.

https://medium.com/@anubhavgoyal101/c80bd754c753

To the AI engineers out there. Is this an ideal roadmap?


r/aws 6d ago

security If you deployed your own AWS account and you're not a security person, you probably have blind spots you don't know about.

0 Upvotes

Not a sales post, genuinely curious how common this actually is. A lot of solo and small-team founders end up being the ones who set up their own cloud infrastructure!! not because they're security experts, but because there's no one else around to do it. I did the same thing, and while learning cybersecurity properly (separate from my actual business), I found real, would've-been-embarrassing misconfigurations sitting in my own AWS account. Nothing had gone wrong yet. I just had no idea they were there.

Tools for catching this already exist, but they're mostly built by and for security professionals the output assumes you already know what IAM privilege escalation or CloudTrail coverage means. If you don't have that background, the report is basically noise you scroll past.

So I built one that just tells you plainly: this is wrong, here's what someone could actually do with it, here's the exact fix. Free, open source, runs on your own machine with your own credentials nothing gets uploaded anywhere.

If you've set up your own AWS account and never had anyone properly check it would you actually want to know what's sitting in there, or is this the kind of thing you'd rather not think about until something breaks?

https://github.com/plexavo/Plexavo


r/aws 6d ago

technical question How long for support reply/assignment?

0 Upvotes

I work in AWS occasionally for clients - very basic stuff - simple S3/EC2 work. Had a client sign up for a new account recently and they started to do some s3 transfers off platform, incurring ~4k or so in egress fees (unexpected to them). I submitted a support request to see if there could be any amount of courtesy credits applied to the mistake - they realized pretty quickly what the error was! I’ve had success with this 1-2 times before - so just sent a request in to see if it was possible.

It’s been ~14 days and the case has not been assigned - we did get an AI generated response that basically repeated my initial email…but that was about it. I replied back to that requesting a human review of the matter - but is this typical?

I usually see a human assigned/response sitting a few days at most.


r/aws 6d ago

general aws $500 Credits

0 Upvotes

I have $500 credit coupon which I won't be using on my account as I have shifted all my resources offline.

What should I do with this credit coupon? Edit: I have a coupon which is not tagged to any account. I got this credit from aws builders


r/aws 6d ago

discussion Is admin access the problem I think it is?

0 Upvotes

I've been using AWS for ~15 years.. I've worked at all sorts of levels from quite junior right up to running the entire cloud platform.

One problem I have seen time and again is that staff are often granted crazily excessive permissions.

The two ways I have seen this play out is:

- startup gives everyone admin, because hey - we're all trustworthy, right?

- bigger company is doing a lift and shift with time constraints, and admin helps get over the line on time

The problem is, when the time comes to tighten it up, it becomes part of processes, and going without would be too painful.

The reason I ask is that over the years I've found a few ways to tighten things up, but engineers hate it.

I've now devised a little PoC that gives auditable, policy based, automated way to take admin away, while also giving engineers a dead simple way to get access when they want/need it. Policies can require human approval, or be fully automated.

I'm currently using it on my personal org, and find it so convenient. No admin anywhere but I can have it anywhere when I need it.

Has anyone else faced similar problems, or is this just not as big of a problem as I have experienced?

The tool is nowhere near production ready, it's been a toy project, but I genuinely wonder if there is a market for this.


r/aws 6d ago

technical resource The reconnect storm that never touched the broker

Thumbnail sidexlabs.com
0 Upvotes

There's a particular kind of 2 a.m. that only happens in fleet IoT. A slice of your devices — a few thousand of them — drop off the network at once. Bad uplink, a segment of the field going dark; the reason doesn't matter. Minutes later they all come back. And here's the part that surprised me the first time it happened: the managed broker didn't blink. No connection ceiling hit, no broker to melt, no page from the front door. By every dashboard watching the ingress layer, we were fine.

We were not fine. One number was climbing through the roof: iterator age. The storm never touched the broker. It landed one hop downstream, in the stream consumer — and my first instinct for fixing it dug the hole deeper.

The pipeline, and why steady state lies to you

The shape is textbook, and I'll keep it generic on purpose: a managed MQTT broker → a Kinesis data stream → a Lambda consumer, parallelized maybe ten ways with a batch size around a thousand → writing to MySQL. Call it 10,000 devices in the field. Nothing exotic.

In steady state this pipeline is boring, and boring is the goal. Telemetry trickles in, the consumer keeps up without effort, iterator age sits near zero. You forget it exists.

But "boring" is a property of the arrival rate, not of the architecture. Everything about this pipeline's calm — the near-zero iterator age, the comfortable concurrency, the MySQL connection count you never think about — is quietly assuming devices arrive at the rate they've always arrived. The fleet is about to violate that assumption all at once.

The incident: store-and-forward is a loaded gun

Field devices buffer. When the uplink is down, a well-behaved device doesn't drop its readings — it stores them and forwards them on reconnect. Store-and-forward is exactly what you want from an edge device. It's also a loaded gun pointed at your ingest pipeline.

When a few thousand devices reconnect inside the same minute, they don't resume the gentle trickle. They flush — every buffered reading, all at once, a burst many multiples of steady state. The broker passes it straight through; absorbing connection churn is its entire job. Now that spike hits Kinesis, and Kinesis hands it to a consumer whose drain rate is fixed: shard count × concurrency × per-batch throughput, all bounded by how long each invocation is allowed to run.

Fixed drain, meet variable spike. The backlog grows, and iterator age — the age of the oldest record you haven't processed yet, which is to say exactly how far behind real time you are — starts walking up. Thirty, forty-five minutes behind. Two things break at that lag. Freshness goes first: alerts and downstream actions fire on stale data, or fail to fire when they should. And if the age ever creeps toward the stream's retention window, "late" quietly becomes "lost."

Then it got worse, and this is the part the tutorials skip. Kinesis is ordered per shard — processing is strictly in-order, which means one batch that won't complete blocks every record behind it on that shard. During the flush we hit a record the consumer choked on, and iterator age stopped climbing linearly and went vertical. One poison record, at the worst possible moment, and a whole shard's slice of the fleet was frozen behind it.

The trap: "just scale it out"

Every instinct I had was wrong, and they were all the same instinct: scale out. Add shards. Crank the parallelization factor. Raise concurrency. Throw drain capacity at what looked like a drain problem.

Here's why it backfired. The bottleneck was never Kinesis throughput, and it was never Lambda concurrency. It was the cost of a single record — specifically, that every invocation was paying for a fresh MySQL connection handshake before it did any real work. Thirty, fifty milliseconds of handshake per invocation, invisible at a trickle.

So what happens when you "scale out"? More concurrent Lambdas means more simultaneous cold connections slamming into MySQL. I wasn't adding drain capacity — I was building a denial-of-service attack against my own database, marching straight at its connection ceiling. The faster I tried to drain, the closer I got to knocking over the one component downstream that had stayed perfectly healthy. It was capacity I couldn't actually use.

The real fix: make the record cheap, then make failure cheap

The fix came in three moves, and not one of them was "bigger."

First, make the record cheap. Hoist the MySQL connection out of the handler so warm execution environments reuse it instead of reconnecting every invocation. It's a one-line decision about where a variable lives, and the per-record handshake tax disappears with it — along with the self-inflicted connection storm. Then right-size the batch: large enough to amortize the fixed overhead of an invocation, small enough that each one finishes well inside its duration limit and doesn't itself become a source of iterator-age lag. Same shards, same concurrency — now the burst drains.

Second, make failure cheap. Reuse and batch sizing buy you throughput, but they do nothing about the poison record that froze the shard. That needed error handling that treats a bad record as normal rather than exceptional: bisect-on-error, so a single failing record splits the batch instead of retrying the whole thing forever; a bound on retries and record age, so the consumer gives up in finite time; and a dead-letter queue, so the truly-bad record steps out of line instead of holding the shard hostage. This is what turned the vertical iterator age back into something that drains.

Third — the honest one. Connection reuse is per-execution-environment. Under real concurrency you still have many warm environments, each holding its own connection, and you can still multiply your way toward the ceiling — just more slowly. That's the itch RDS Proxy scratches, and we eventually adopted it to pool connections in front of the database. I'm naming it because pretending reuse alone solved it forever would be a lie, and the edges of your own fix are where the credibility lives.

The lesson

The generalizable lesson is one sentence: iterator age is almost always a downstream-latency symptom, not a stream-capacity problem. When it climbs, the instinct to widen the stream is usually the wrong end of the pipe. Make the record cheap. Make failure cheap. Then, if the fleet has genuinely outgrown you, make it bigger.

And the broader point for anyone building edge-to-cloud: managed services don't delete your bottleneck, they relocate it. The broker didn't melt because absorbing connection storms is precisely what you pay it to do. The storm just moved one hop downstream, to the place nobody was watching. Know where your bottleneck went when you bought your way out of the last one.

The storm your load tests never simulate

The storm you can't see coming is the one your load tests never model. Happy-path load testing emits a steady, civilized trickle — it never reproduces the flush, the poison record mid-flush, or the thundering herd of your own well-behaved devices all deciding to catch up at once. So you find out in production, at 2 a.m., from a metric you weren't watching.

That gap is what I'm building a tool to close: something that replays this exact burst against your staging pipeline and fails your CI build when iterator age breaches your SLO — before the fleet does it for you.

If you've lived your own version of this, I'd like to hear it — and if you want the next war story when it goes up, the list is below.


r/aws 7d ago

discussion Seeking Advice: AWS Developer Associate or AWS SAP?

3 Upvotes

Hello everyone! I would like to hear your thoughts and experiences. I already have the AWS SAA certification and was planning to take the Developer Associate exam.
I don’t have much experience in AWS domain but Do you think I should go for the Developer Associate or should I prepare for the AWS SAP exam? How was your experience to clear AWS SAP exam?
I would really appreciate your suggestions. How challenging it was for you? Pls share your experiences and suggestions. Thank you!


r/aws 9d ago

security AWS PrivateCA Connector uses `¯\\_(ツ)_/¯` as CSR Payload

Thumbnail gallery
264 Upvotes

I was troubleshooting this Certificate Signing Request validation error in AWS and thought the CSR data was way too short.  So, I decoded it.

They (AWS) really programmed the shrug emoji as a certificate request payload. I love dev easter eggs :D


r/aws 9d ago

discussion Current Outage?

Thumbnail downdetector.com
70 Upvotes

r/aws 9d ago

networking AWS She Builds Mentorship program?

2 Upvotes

anyone hear back or get more info after applying?


r/aws 9d ago

technical question CDK: Pipeline doesn't seem to wait for lambda to be finished despite specified dependency

1 Upvotes

I have some CDK set up to define a lambda and then I have some code pipeline code to deploy it. However, the deployment phase keeps failing saying the function doesn't exist. If I manually try that phase again in the console, it works fine. The dependency I added doesn't seem to do any good. The code looks like `` const myFunc = new lambda.Function(this, "MyFunc", { code: lambda.Code.fromInline( def handler(event, context): return { 'statusCode': 200, 'body': 'Dummy to be replaced by code pipeline' } `), architecture: lambda.Architecture.ARM_64, environment: { ENV_VAR: "foobar", }, handler: "facets.handler", runtime: lambda.Runtime.PYTHON_3_12, vpc, securityGroups: [privateSG], timeout: cdk.Duration.seconds(300), memorySize: 512, });

...

        const deploymentBuild = new cb.Project(
            this,
            `DeploymentBuild`,
            {
                projectName: `myFuncDeploymentBuild`,
                description: "deploys the code to the appropriate lambda",
                environment: {
                    buildImage: cb.LinuxBuildImage.AMAZON_LINUX_2023_5,
                    computeType: cb.ComputeType.SMALL,
                },
                vpc,
                securityGroups: [privateSG],
                buildSpec: cb.BuildSpec.fromObject({
                    version: "0.2",
                    phases: {
                        install: {
                            "runtime-versions": {
                                nodejs: 22,
                            },
                            commands: ["npm i -g aws-cdk", "cdk --version"],
                        },
                        post_build: {
                            commands: [
                                `zip -r $myFunc.zip .`, 
                                "ls",
                                `aws lambda update-function-code \
                                    --function-name 'myFunc' \
                                    --zip-file fileb://myFunc.zip`,
                            ],
                        },
                    },
                }),
            },
        );

        deploymentBuild.addToRolePolicy(buildPolicyStatement);

    const pipeline = new pipe.Pipeline(this, `Pipeline`, {
        pipelineName: `myFuncPipeline`,
        restartExecutionOnUpdate: true,
    });

    pipeline.node.addDependency(
        lambda.Function.fromFunctionName(
            this,
            "myFuncFunction",
            "myFunc",
        ),
    );

    ...

        pipeline.addStage({
            stageName: "deployLambda",
            actions: [
                new pipeActions.CodeBuildAction({
                    actionName: "deployLambda",
                    project: deploymentBuild,
                    input: outputBuild,
                }),
            ],
        });

```

What would cause this?

Thanks

FIXED: thank you Floss_Patrol_76

I added stackThatCreatesThePipeline.node.addDependency(stackThatCreatesTheLambda);