Back to the registry
Agent passport

Nvidia Nemotron-3-Super-120B-FP8 Stained Glass Enabled

Protopia AI, Inc. · Cybersecurity & IT

No attestation published

Certification per AWS Marketplace.

Provenance reach4 of 12 layers traced

Evidence tier Source Confirmed · 4 captures on record

User ratingNot rated0 reviews on the listing
Runs onUnknownML model
ProvenanceUnknown44% of the provenance layers this product can disclose
Evidence riskHighSign in to see the basis for this band.

What the publisher says

As described on AWS Marketplace.

# Roundtrip Protection for Sensitive AI Workloads

Organizations in many sectors (including, but not limited to healthcare, financial services, and legal) need to run large language models over confidential data - but shared GPU infrastructure and multi-tenant environments create exposure risk. Protopia AI's Stained Glass Transform (SGT, not included in this marketplace product) protects model inputs through stochastic transformation. This listing is a companion model endpoint configured to accept SGT-protected representations as well as plaintext, and return protected responses with Stained Glass Output Protection.

Show the rest of the publisher’s description (34 more lines)

Protopia AI's Stained Glass Output Protection encrypts tokens as they are generated by an LLM. The generated response is never fully available in plaintext on the compute infrastructure. Raw data never leaves its trust zone-only these protected re-representations of the data reach the compute infrastructure.

## How It Works

Stained Glass Output Protection encrypts each token as it is generated by the LLM. Every request establishes a secure key exchange to individually encrypt each generated token. The plaintext response never exists in its entirety on the compute infrastructure. This prevents leakage via logs, dumps, state, etc. Only the authorized client holding its private key can decrypt the response.

Not included in this marketplace product, the companion Stained Glass Transform (SGT) transforms your input data into unintelligible prompt embeddings (prompt_embeds) before being sent to the compute infrastructure. The plaintext never leaves the data owner's trust zone. This [SGT-Enabled Nvidia Nemotron-3-Super-120B-FP8] supports SGT-transformed inputs, processing protected representations directly with no decoder on the endpoint.

This roundtrip protection pipeline means that neither the input nor the output is ever exposed on the inference host - a critical requirement for regulated industries handling patient records, financial

documents, or privileged legal communications.

## Use Case: Confidential Document Summarization

A financial services firm running retrieval-augmented generation (RAG) over proprietary research documents on shared cloud GPUs can deploy this listing to ensure that neither the retrieved context

(input) nor the generated summaries (output) are visible to the infrastructure operator or co-tenants.

The same pattern applies to healthcare organizations summarizing patient records or legal teams

analyzing privileged communications.

## About the Model

The underlying NVIDIA Nemotron-3-Super-120B-A12B model was pre-trained on over 25 trillion tokens

spanning code, math, science, and general knowledge across 20 languages, then instruction-tuned for

tool calling, structured outputs, and long-context retrieval. NVIDIA reports benchmark results

including 86.01 on MMLU and 90.67 on GSM8K (8-shot). This listing serves an FP8-quantized variant for

reduced memory footprint and inference latency, which may show minor accuracy variation from published

BF16 figures.

## Deployment and Scaling

Served on vLLM with tensor parallelism that auto-sizes to the deployed instance's GPU count.

Verified serving prompt_embeds correctly across all supported real-time instance sizes.

Note: despite the delivery configuration listing a recommended batch transform instance, batch (asynchronous) inference is not currently supported -- no AWS SageMaker batch transform instance type is capable of serving a model this large. Only real-time inference is supported today.

## Why Stained Glass Transform Over Alternatives

Unlike homomorphic encryption, SGT operates with low compute overhead on commodity hardware - no specialized TEE hardware required. Unlike differential privacy approaches that inject noise, SGT

preserves full data fidelity so model accuracy is maintained. This frees powerful GPUs for inference rather than encryption overhead, and integrates into existing AI pipelines with minimal latency impact.

## Getting Started

To evaluate this solution or schedule a technical demo with our solutions engineering team, contact us

at contact@protopia.ai. We support deployments across on-premises, multi-tenant cloud, and edge

environments.

## Input/Output Format

Inputs can be passed as prompt_embeds (transformed embeddings). Outputs are returned as tokenwise-encrypted

payloads using a secure key exchange. See the sample Jupyter notebook

(https://github.com/protopia-ai/sagemaker-marketplace/blob/main/invoke_sgt_nemotron_endpoint.ipynb)

for a full walkthrough of payload structure, decryption steps, and SDK usage.

Highlights

Highlighted by the publisher on AWS Marketplace.

Unlock AI Data Potential Securely Access sensitive data without compromising security, enhancing AI model accuracy with high-quality inputs on shared infrastructure.

Protect the output, too Generated tokens are returned tokenwise-encrypted.

NVIDIA Nemotron-3-Super-120B (FP8); a large, high-quality open model, now with full input+output protection.

Agent build and provenance

See the full provenance

The layer-by-layer build, the evidence behind each claim, the risk basis and the cross-marketplace links are open to any account. Some rows are disclosed, some the source leaves Unknown; a free account shows you which.

Compliance

Government
  • FedRAMPConfirmedNot listed90%, registry-checkedNo FedRAMP Marketplace entry matched this vendor's domain, checked 2026-08-27registry recordas observed 2026-08-27

Confirmed means matched to a public authoritative registry. Claimed means the vendor or its listing states it, not yet cross-checked. A framework not shown was not found in any source we hold, which is not evidence against it. Not listed means a scoped registry check found no match for this vendor's domain: a No is a scoped registry check, not a compliance judgment. Confidence bands: 95% domain-verified, 90% registry-checked, 80% self-attested, 70% weak signal. Self-attested items marked “vendor's site” are gathered from the vendor's own website and are not verified by us.

Plans and pricing as listed

6 listed
ml.g5.48xlarge Inference (Batch)
  • HostHrs
$0.00
ml.g7e.48xlarge Inference (Real-Time)
  • HostHrs
$0.00
ml.g7e.12xlarge Inference (Real-Time)
  • HostHrs
$0.00
ml.g7e.24xlarge Inference (Real-Time)
  • HostHrs
$0.00
ml.g6e.12xlarge Inference (Real-Time)
  • HostHrs
$0.00
ml.g6e.48xlarge Inference (Real-Time)
  • HostHrs
$0.00

Refund terms

As stated by the publisher on AWS Marketplace.

This package is provided free of charge, and no refunds will be provided.

Sources

Marketplace listingaws.amazon.comSource
App certificationaws.amazon.comSource
StandardEulaStandardEulaSource

Publisher resources

1 link
vLLM Prompt Embeddingsdocs.vllm.aiSource

Linked repositories

1 repo
protopia-ai/sagemaker-marketplacegithub · AWS MarketplaceSource

Unknown means this listing does not publish a repository. It is not a statement that the code is closed, and a linked repository is not a claim that the publisher wrote it: the registry computes that relationship privately and does not publish it.

Pricing
Paid
6 plans listed
Delivery
ML model
Please contact us for further support at support@protopia.ai
Open the source listing ↗

Evidence risk is the share of the build you cannot see before you deploy, not a security rating. Sign in to see the layer-by-layer basis for this band.