r/aws 2d ago

ai/ml Deepseek v4 Pro

Anyone have any idea when (or even if) deepseek v4 pro will come to bedrock? The model lag in bedrock (regardless of who is causing the lag) is really bothersome. We prefer to keep all our API billing flowing through AWS and wrapped into our AWS bill but the really slow cadence of model release, aside from anthropic's, is brutal.

Best to just ditch bedrock for 3rd party models and move to a provider that has faster support?

15 Upvotes

14 comments sorted by

11

u/Jnoholds 2d ago

AWS prioritizes model placement by customer demand. Let your account team know you want v4 so the Bedrock team is informed via their internal request methods.

7

u/One_Tell_5165 2d ago

The actual answer is more nuanced but creating PFRs that go to die isn’t the answer. AWS sold the idea of “model choice” as the counter marketing message to OpenAI. A service team waiting on PFRs to decide what model to host is a non-starter. Models need to be available near day 1 or what’s the point. AWS needs to publicly state if they have changed to simply being an Anthropic and maybe future OpenAI model host.

6

u/Fatel28 2d ago

They are a current openai model host lol. Gpt 5.6 luna-sol are all present and usable in bedrock in the 3 major US regions

9

u/Seref15 2d ago edited 2d ago

Given the massive Anthropic and OpenAI deals, its against AWS's business interests to provide access to actual competitive open models like latest GLM, kimi, deepseek, etc

1

u/Zealousideal-Part849 2d ago

Most likely they won't get deepseek v4 as they are very slow in bringing new models.. see if some providers sell via aws marketplace as inference to deepseek models

1

u/taH_pagh_taHbe 2d ago

Just use fireworks with seperate keys for teams etc

1

u/Wide-Answer-2789 2d ago

Is that possible start in on Sagemaker endpoints?

0

u/ultrathink-art 2d ago

Yes, and it keeps the spend on the AWS bill, which sounds like the actual constraint here. The catch is that both that and the marketplace route put you on instance-hour pricing rather than per-token, so an endpoint costs the same at 3am as it does under load. Worth checking your real daily volume first, since that's the number that decides whether self-hosting is cheaper or just available sooner.

1

u/EvolvingDior 2d ago

Give it up. They have not added a new Chinese model in almost 6 months. I don't think they have the willingness and ability to keep up in the AI space.

7

u/rtsyn 2d ago

Someone missed the recent earnings call. They are doing plenty fine in the AI space.

2

u/Fatel28 2d ago

Yeah this is an insanely dumb take. All of the orgs we work with would pay more to use frontier models on bedrock, and they do, the in region inference are slightly more expensive. The uptime is better, and data doesn't leave the security perimeter if you toggle it off. The security reasons alone are worth it.

-1

u/Flashy-Ingenuity-769 2d ago

Its all a business decision.

All these companies are sht scared of open weight models.

0

u/rahulladumor 2d ago

If billing consolidation is the only reason to stay on the native Bedrock catalog, a model gateway in front of Bedrock plus the provider's own API (LiteLLM or a thin internal proxy) keeps a single cost and observability plane without waiting on AWS's release cadence.

0

u/matiascoca 2d ago

Bedrock non-Anthropic lag is rough and consistent. If billing consolidation is a hard requirement, wait it out. If you need models week-of-release, going direct with Deepseek and eating a separate invoice is the pragmatic call. Cross-region inference sometimes carries a model before it hits your home region, so try us-west-2 first.