A natural gas compressor station and pipeline manifold at sunset
Frequently asked questions

HillStar for RFP Evaluation

Frequently asked questions for procurement teams in regulated industries.

Overview

What does HillStar do?

HillStar reads every submission you receive and evaluates it against your criteria — and shows you the source text behind every conclusion it reaches.

You provide the solicitation and the responses. HillStar returns each respondent scored criterion by criterion. Every score carries its written rationale and the passage from that respondent's own document that supports it: file, page, and quoted text, one click from the score.

Nothing is asserted without a source. Where a response doesn't support a finding, HillStar reports the absence and cites where it looked — it does not fill the gap with inference.

Your committee reviews the analysis, forms its own conclusions, and makes the award recommendation.

Why use HillStar instead of a general-purpose AI tool or manual review?

Procurement teams in regulated industries deserve more than a generic AI chat tool or another spreadsheet. HillStar is built around the real evaluation workflow: the solicitation sets the standard, every respondent is measured consistently, the committee keeps its judgment, and the final record must withstand scrutiny.

HillStar completes the first-pass reading, comparison, and evidence collection across every submission. Every criterion is grounded in the solicitation, and every finding links to the exact source text, file, and page. Your committee can then focus on the work only it can do — weighing findings, applying judgment, and making the award recommendation.

It does not replace your process or make the award decision. It gives your team a professional, purpose-built system for producing the cited, auditable record behind it.

What changes when we use HillStar?

Your evaluation methodology is published in the solicitation. Your committee's scores become the record. If the award is challenged, every number has to trace back to something a respondent actually wrote.

None of that changes.

What changes is that the analysis is already done when your committee sits down. Every response, every criterion — scored against your standards, with a written rationale and the source page behind each finding, in a single unattended run. Your committee spends its time weighing findings and forming determinations instead of spending the calendar assembling them.

What problem does it solve?

Complete analysis, at full depth, no matter how large the stack is.

Evaluation depth is normally rationed by page count against a closing date — a dozen responses of several hundred pages each, a pricing workbook per respondent, a committee of people with day jobs. Something gets less attention, and it's usually the tail of the stack: the fourth respondent's exceptions list, the sub-consultant qualifications, the requirement answered in an appendix instead of where the response template asked for it.

HillStar performs the complete first-pass evaluation. Every criterion, for every respondent, scored with a written rationale and a citation to that respondent's own language — the twelfth response analyzed to the same standard as the first.

Your committee then does the work only it can do: weigh the findings, apply the judgment the solicitation reserves to it, and make the determination.

Who is it for?

Regulated buyers — utilities and energy companies, including publicly-owned municipal utilities — where an award has to withstand a protest, a records request, and an audit.

Is anyone using this today?

Yes. HillStar is in production at regulated utilities, running live solicitations.

Our customers' procurement records are confidential and utilities generally don't want their sourcing activity publicized, so we don't publish names or logos without written permission. On a call we can describe the categories and solicitation types we're running, and where a customer has agreed, arrange a reference conversation with a peer in your industry.

Does HillStar make the award decision?

No. The evaluation of record and the award recommendation are your committee's.

HillStar produces analysis and evidence. It applies your criteria uniformly to every response and shows the basis for every conclusion. Your committee weighs that, applies the judgment the solicitation reserves to it, and documents its determination through your normal process.

We're deliberate about this boundary. An award recommendation that originates with a vendor's software is not a recommendation your committee can defend as its own.

How long does an evaluation take?

Minutes. A single run can evaluate dozens of proposals and thousands of pages.

One utility's evaluation covered 22 vendors and roughly 19,800 pages — work conventionally scoped at six to nine months — and came back complete and cited in nine minutes.

The speed is the point, and it compounds. Because a run costs minutes rather than another review cycle, you can re-score after an addendum or a revised submission and still move the award while your pricing is current.

Scoring methodology

Where do the criteria come from?

From your evaluation framework, applied to the specific solicitation.

Your framework lives in HillStar as a category library: the categories you evaluate, and the measurement guidelines beneath each. For a given RFP, HillStar derives the criteria by matching those guidelines against what the solicitation actually requires.

Every criterion must cite the provision in your RFP or SOW that produced it. That requirement is enforced — a criterion HillStar cannot ground in your own solicitation documents is rejected before you see it. Respondents are never measured against a standard you didn't publish.

How does it improve scoring consistency?

The same criteria, the same scale, and the same evidentiary standard are applied to every response.

Consistency is hard to hold across any committee scoring exercise — two people reading the same response against the same criterion can reach different numbers. HillStar's analysis doesn't vary by reader or reading order. Your committee still exercises judgment, but it does so from a consistent baseline rather than reconciling divergent readings.

What does a score come with?

Every score comes with a written explanation and the exact text that backs it up — quoted from the respondent's own document, with the file and page number.

Each piece of evidence is marked HIGH, MEDIUM, or LOW. That tells you the difference between something a respondent clearly committed to in writing and something read from nearby text. If the support is weak or missing, HillStar says so and scores it that way instead of writing a reason to fit.

How do we know the AI isn't inventing findings?

Because every claim it makes is anchored to a document you can open, and the ones that aren't get removed rather than shown to you.

Three things enforce that:

Criteria are grounded in your solicitation. Every criterion must cite the provision in your RFP or SOW that produced it. A criterion that can't be traced to your own documents is rejected before it reaches you — so respondents are never scored against a standard you didn't publish.

Findings are grounded in the respondent's own words. Every conclusion about a response carries the quoted passage supporting it, with file and page. Click it and the respondent's text opens in context, so you can confirm nothing was read out of context or stitched together from unrelated sections.

Citations are validated against the actual document. A citation that doesn't resolve to real text in the real file is discarded, not displayed. You are not asked to trust a page reference — the quoted text is retrieved from the document itself.

How are mandatory requirements and certifications handled?

They become scored criteria, tied to the provision that imposed them, and absence is reported explicitly.

“The response states it is not currently ISO 9001 certified” is a cited finding rather than an empty cell. You see the same requirement answered across all respondents side by side, each with its page and quoted text — the check that surfaces an unsigned certification in the low bid during evaluation rather than during a debriefing.

To be precise about scope: HillStar does not make responsiveness determinations and does not disqualify respondents. Every requirement is scored and evidenced. Determinations of responsiveness and responsibility remain with your committee, where your procurement rules place them.

What if we disagree with a score?

Then you have the citation in front of you and can settle it in seconds rather than debating impressions.

If the evidence is right but the weight given to it is wrong, that's your committee's call to make and document — HillStar's analysis is an input to your evaluation, not a substitute for it. If the criterion itself was too blunt, sharpen it or the underlying guideline and rescore. HillStar creates a new evaluation alongside the prior one; both stay on the record, so the change is visible rather than silent.

What you can't do is edit a score in place, and that's deliberate. An analysis record that can be quietly adjusted after the fact is worth nothing in a hearing.

The evaluation record

What does the committee actually receive?

  • All respondents compared side by side across every category
  • A scorecard per respondent: per-criterion reasoning, evidence, confidence labels, and citations that open to the cited page
  • Every solicitation question answered by every respondent, side by side, each with its citation
  • Complete version history of the rubric and of every scoring run

How does this hold up in a protest?

The record can't be revised after the fact — by your team or by us.

Rubric versions are append-only, enforced at the database level rather than by policy. Rescoring creates a new evaluation instead of overwriting the previous one. Every prior scorecard remains open for inspection with its author, timestamp, and change comment.

So when a question arrives eighteen months later about why a respondent scored a 3 on safety, the criteria in force at that time, the score, the rationale, and the quoted source text are all retrievable as they stood. That's also why HillStar's scores aren't editable in place: an analysis record that can be quietly adjusted has no evidentiary value.

Does it support debriefings?

That's one of the more common uses. Each criterion carries a written rationale and a citation to the respondent's own language, which is the material a substantive debriefing requires — you can tell an unsuccessful respondent specifically where their response fell short and quote what they submitted.

Can we get the record out for our files or a records request?

Yes. On request we produce a complete structured extract: every record in your workspace, down to criterion-level scores and the audit log, together with every document you uploaded. Individual documents your team can download directly at any time.

Plainly: one-click export of a full evaluation is on the roadmap and not yet shipped. Today that extract is produced by us on request, which is fast but is not self-service.

Are we obliged to disclose to respondents that AI was used?

That's a policy and legal question for your organization, not something we'll advise you on — and the answer varies by jurisdiction, by your procurement code, and by what your solicitation already says.

What we can do is give your counsel and your IT reviewer a straight account of how the analysis is produced, what data leaves your environment, where it's processed, and what the audit record contains. Several of our conversations start with exactly that review. If you need specific language for a solicitation or an evaluation plan, we'll support it with facts rather than draft it for you.

Pricing analysis

How is pricing compared?

On a normalized per-unit basis, with provenance to the source cell.

Where a pricing workbook forms part of your solicitation, HillStar reads each respondent's returned workbook and aligns the bid lines, normalizing unit-of-measure differences so the comparison is valid. You get medians, low-bid comparisons, and traceability back to the exact cell in the respondent's own spreadsheet.

Where a respondent restructures your schedule or bids an alternate, HillStar flags the deviation rather than reconciling it. A silent reconciliation is how a price comparison becomes indefensible; that call belongs to a person.

Pricing analysis is being rolled out progressively and may not yet be enabled on your account.

Data handling and security

Where is our data stored and processed?

In the United States, in full.

Your records, uploaded documents, backups, and usage data are stored and processed in US Google Cloud regions and served from a US region, encrypted at rest and in transit. The database has no public IP address, and storage enforces public-access prevention.

AI processing runs on US-based commercial providers. Our OpenAI traffic is pinned in code to OpenAI's US endpoint with no fallback — a non-US credential fails rather than silently succeeding. Embedding traffic is restricted in code to US-only regions and fails closed. Anthropic, which performs the primary scoring work, processes in the United States.

Is our data used to train AI models?

No. Your content is transmitted to our providers to evaluate your solicitation and for no other purpose. Under the commercial API terms we operate on, Anthropic and OpenAI do not use API customer content to train their models.

Are you SOC 2 certified?

SOC 2 is in progress and expected Q3 2026. See our Trust Center for detailed information about status and policies.

Google Cloud, our infrastructure provider, holds independent SOC 1/2/3, ISO 27001, and FedRAMP High authorization.

Who can access our evaluation?

Only members of your workspace.

We enforce workspace permissions with Admin, Member, and Viewer roles, with full traceability.

Implementation

Do we have to change our procurement process?

No — not your documents, your standards, or your workflow.

HillStar takes the package your process already produces: solicitation, SOW, technical specifications, question set, pricing schedule, and responses, exactly as issued and received. Files are classified on upload; nothing needs restructuring or re-tagging. Your evaluation standards become your category library rather than being replaced by someone else's.

On footprint, straightforwardly: HillStar maintains its own record of the evaluation. It does not replace your procurement system and does not integrate with it — documents are uploaded rather than transferred by interface.

What formats do you accept?

PDF, DOCX, XLSX, CSV, TXT, and Markdown. Up to 35 files per upload, 100 MB per file. Scanned and image-only PDFs are processed through OCR automatically where text cannot be extracted directly.

PowerPoint, image files, and legacy .doc / .xls are not accepted. If your respondents typically submit something outside that list, send a sample before your closing date rather than after.

See HillStar evaluate your next RFP

Bring a real bid stack. We'll show you scored, ranked, evidence-backed results in a single working session.