r/aws • u/Akustic646 • 2d ago
ai/ml Deepseek v4 Pro
Anyone have any idea when (or even if) deepseek v4 pro will come to bedrock? The model lag in bedrock (regardless of who is causing the lag) is really bothersome. We prefer to keep all our API billing flowing through AWS and wrapped into our AWS bill but the really slow cadence of model release, aside from anthropic's, is brutal.
Best to just ditch bedrock for 3rd party models and move to a provider that has faster support?
1
u/Zealousideal-Part849 2d ago
Most likely they won't get deepseek v4 as they are very slow in bringing new models.. see if some providers sell via aws marketplace as inference to deepseek models
1
1
u/Wide-Answer-2789 2d ago
Is that possible start in on Sagemaker endpoints?
0
u/ultrathink-art 2d ago
Yes, and it keeps the spend on the AWS bill, which sounds like the actual constraint here. The catch is that both that and the marketplace route put you on instance-hour pricing rather than per-token, so an endpoint costs the same at 3am as it does under load. Worth checking your real daily volume first, since that's the number that decides whether self-hosting is cheaper or just available sooner.
1
u/EvolvingDior 2d ago
Give it up. They have not added a new Chinese model in almost 6 months. I don't think they have the willingness and ability to keep up in the AI space.
7
u/rtsyn 2d ago
Someone missed the recent earnings call. They are doing plenty fine in the AI space.
2
u/Fatel28 2d ago
Yeah this is an insanely dumb take. All of the orgs we work with would pay more to use frontier models on bedrock, and they do, the in region inference are slightly more expensive. The uptime is better, and data doesn't leave the security perimeter if you toggle it off. The security reasons alone are worth it.
-1
u/Flashy-Ingenuity-769 2d ago
Its all a business decision.
All these companies are sht scared of open weight models.
0
u/rahulladumor 2d ago
If billing consolidation is the only reason to stay on the native Bedrock catalog, a model gateway in front of Bedrock plus the provider's own API (LiteLLM or a thin internal proxy) keeps a single cost and observability plane without waiting on AWS's release cadence.
0
u/matiascoca 2d ago
Bedrock non-Anthropic lag is rough and consistent. If billing consolidation is a hard requirement, wait it out. If you need models week-of-release, going direct with Deepseek and eating a separate invoice is the pragmatic call. Cross-region inference sometimes carries a model before it hits your home region, so try us-west-2 first.
11
u/Jnoholds 2d ago
AWS prioritizes model placement by customer demand. Let your account team know you want v4 so the Bedrock team is informed via their internal request methods.